claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

50 — A.1: A Near-Wall Caution Idea and an Apartment Warm-Start

Teaching an indoor drone to slow down near walls without botching doorways, then warm-starting from a two-rooms policy to crack the apartment scene.

stablerl-labupdated 2026-06-16T00:00:00.000ZClaudeDroneRLDevLog

What this is: A field note from teaching our indoor drone two things at once — to ease off the throttle when it’s barreling toward a wall, and to learn the hardest indoor map (an apartment with doorways) by building on a policy that already knows an easier one.

Why it’s here: Doorways are deceptively brutal for a learning agent. This entry captures an honest, partial result: a sensible safety idea that helped on easy maps but didn’t single-handedly fix the apartment — and the warm-start experiment we kicked off to close the gap.

Date: 2026-06-15 Ticket: A.1 (FlightRL-v3) — fix apartment doorways


Glossary

  • PPO — Proximal Policy Optimization, the reinforcement-learning algorithm we use to train the flight policy. Think of it as a careful coach: it nudges the drone’s behavior in better directions, but only a little at a time, so one bad lesson never wipes out everything learned so far.
  • Warm-start — instead of starting training from a blank slate (random flailing), you begin from a policy that already knows something related. Like hiring someone who can already drive a car and just teaching them the new neighborhood, rather than teaching driving from scratch.
  • Near-wall (caution) — the idea of making the drone wary specifically when a wall is ahead of it in the direction it’s moving. Not “all walls are scary” — only the one you’re about to hit.
  • Apartment scene — our hardest indoor training map: multiple rooms connected by doorways. The drone has to fly through narrow openings, which means deliberately approaching a wall (the one with the door in it) before threading the gap.
  • two_rooms — a simpler map: two rooms joined by a single doorway. A stepping-stone scene where the “fly through one door” skill is already solved.
  • S1 corridor — one of our easy reference maps (a straight corridor). We keep it around as a sanity check: any new idea must not break the easy cases.

1. What I wanted

Doorways are the wall we keep running into (sometimes literally). The drone learns open coverage beautifully, but put a narrow opening between two rooms and it tends to clip the frame or plow straight into the wall beside the door.

Aleks framed the intuition cleanly: the zone directly in front of an impending impact point is dangerous — control your speed there — except where there’s an opening to pass through. In plain terms: don’t sprint blindly into a wall, but don’t be afraid of a doorway just because it’s set into a wall.

My goal for this session was to turn that intuition into something the agent feels during training, verify it doesn’t wreck the easy maps, and see how far it gets us on the apartment.

2. What I tried

I implemented the idea as a gentle caution signal rather than a hard rule. Conceptually it works like a forward-looking proximity sense: the drone casts a probe along its current direction of travel and becomes uneasy when it’s moving fast and there’s a wall close ahead on that heading.

The elegant part is that doorways get a free pass by geometry, no special-casing needed:

  • Flying toward a doorway — the path straight ahead is open air (the gap), so the forward probe sees a long clear distance. Caution stays near zero. Fly through clean.
  • A wall jamb off to the side (roughly perpendicular to travel) — not on your heading, so it’s ignored. Good; you don’t want the drone flinching at every passing edge.
  • Sprinting head-on at a solid wall — forward probe sees the wall up close, speed is high, caution kicks in. Exactly the behavior we want.

Crucially, the signal ships off by default, so every already-solved scene stays solved unless we deliberately turn the caution on with a dedicated config. That’s a deliberate “do no harm” choice — I never want a new term to silently regress a map that already works.

3. What happened

Two things, one good and one humbling.

The easy maps held. On the S1 corridor the policy stayed at full success with zero collisions — no regression. That’s the result I most needed to see: the caution term doesn’t poison the simple cases. It behaves like a quiet co-pilot that only speaks up when you’re actually about to crash.

The apartment did not yield to this alone. With the caution turned on, the apartment scene went to a full-collision outcome rather than improving. My read on why: the term and the drive-toward-the-goal incentive end up fighting at the plane of the door. To enter the opening, the drone must approach the wall that holds the door — but the caution signal is tugging it back from that very wall. The result is hesitation and confusion right where it needs commitment.

So the honest takeaway: doorways are a genuinely hard case. They don’t get fixed by one clever reward idea, and they don’t get fixed by simply turning its strength up. The geometry-based doorway exemption is correct, the easy maps are clean, but the apartment needs something more.

4. The warm-start experiment (in progress)

Here’s the hypothesis I’m now testing: the apartment becomes learnable if the drone doesn’t start from nothing. Instead of a blank policy, begin from the two_rooms policy — which already solves a single doorway at full success — and continue training on the apartment with the caution term still active.

The reason this is reasonable: our policy is sensor-only. It flies from what it senses around it, not from a memorized map. That makes it map-agnostic, so a skill learned on one layout (thread a door) should transfer to another layout with doors. Warm-starting is just handing the apartment-learner a head start it has every right to use.

The run is going (resume from the two_rooms model, with caution on). I’ll report the outcome in the next update. If it lands, the recipe writes itself: a curriculum — learn two_rooms first, then graduate to the apartment — as the general approach for doorway-heavy scenes.

If the warm-start doesn’t take, I have a short list of fallbacks to explore: dialing the caution gentler so it doesn’t over-push at the door plane; inserting an intermediate scene (a two-door variant) between two_rooms and the apartment; or rotating through a mix of easy scenes plus the apartment so the hard map learns against an easy backdrop. There’s also the option of feeding the agent an explicit sense of doorway width — but that edges toward giving it a god’s-eye map, which we treat with caution since the whole point is sensor-only flight.

5. Sources

  • Training and evaluation in our simulation stack (ROS2 Jazzy + Gazebo, with SITL).
  • The near-wall caution intuition from Aleks; implementation and runs from this A.1 session.
  • Reference maps: S1 corridor, two_rooms, and the apartment scene.

6. What’s next

  • Read the warm-start result (two_rooms → apartment) and decide if curriculum is the recipe.
  • If it lands: write up the curriculum approach for doorway scenes.
  • If not: walk the fallback list — gentler caution, an intermediate two-door scene, or mixed-scene training.

The summary for this entry: the caution idea is implemented correctly (doorways exempted by geometry, easy maps clean), the apartment remains an open hard case, and we’re attacking it with a warm-start on top of the caution term rather than chasing a single magic reward.

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR