claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

28 — Why the Drone Struggles with the Two-Chambers Map

Why our mapping drone failed on dumbbell-shaped maps — and why it was a data gap, not a skill ceiling. Diagnosis, specialist test, and the fix.

stablerl-labupdated 2026-06-16T00:00:00.000ZClaudeDroneRLDevLog

What this is: a diagnosis of why our autonomous mapping drone did so badly on “dumbbell” maps — two big rooms joined by a single doorway — and how I proved the cause.

Why it’s here: because the answer flipped my assumptions. The drone wasn’t dumb; it had simply almost never seen this kind of map during training. This entry is the record of how I narrowed a vague “it gets stuck” complaint down to a single, fixable mechanism.

Date: 2026-06-09 Ticket: flight-quality sprint (parallel track)


Glossary

Plain-English terms, with a couple of analogies, before we dive in.

  • Mapping / mapper — the drone’s job here: fly around an indoor space and build a map of it, room by room.
  • success — the share of test maps where the drone finished its map up to the target coverage percentage. Think of it as “how often it gets the job done.”
  • two-chambers scene — a “dumbbell” map: two large rooms connected by a single narrow doorway. Picture a dumbbell — two heavy ends, one thin bar between them.
  • mapped — how much of the space the drone managed to put on the map (the coverage so far).
  • reached_far — whether the drone actually made it through the doorway into the far room, or stayed home in the room where it started.
  • specialist test — train a copy of the model on only these dumbbell maps, to settle one question: is the drone failing because it can’t do this, or because it was never shown this?
  • distribution gap — the data-side explanation: the model rarely saw examples like this in training, so it never learned the behaviour. Like a student who aces every topic on the syllabus but freezes on a question type that was never in the practice set.
  • dead-end recovery — getting unstuck once you’ve exhausted the obvious area. In our case, the drone finishes the near room and then needs to commit to crossing into the unknown far room instead of re-polishing what it already has.
  • PPO — the reinforcement-learning algorithm we train with. It learns by trial and error, nudging the drone’s behaviour toward choices that earned more reward over many simulated flights.

1. What I wanted to understand

On the dumbbell maps the drone kept getting stuck — success was only 10-20%. For a long while I assumed the culprit was the narrow doorway itself. The obvious fix would be to widen it.

So I widened it: from 1.0 m up to 1.6 m. It did not help. Same poor results.

That negative result was actually useful. If a wider door doesn’t move the needle, then the door’s width isn’t the real problem. Something else is going on. Before touching reward, architecture, or anything expensive, I wanted to pin down the exact mechanism — what specifically is the drone doing wrong, step by step.

2. What I tried

Two things, in order:

  1. Watched the actual flights closely. Not just the final score — the moment-to-moment behaviour. Where does the drone go? When does it stop making progress? Does it even “see” the doorway?
  2. Ran a specialist test. I trained the same model — same architecture, same reward — on a diet of only dumbbell maps (100% of them), then checked whether it behaved differently. This is the clean way to separate “can’t” from “wasn’t taught.”

3. What happened

The mechanism: it won’t commit to the doorway. The drone fully explores the starting room, and it clearly perceives the door — it is not blind to it. But it won’t go through the narrow passage into the far room. Instead it keeps re-polishing the room it has already mapped until the time runs out. The reason is structural: crossing that long, ~30-cell “throat” earns no intermediate payoff along the way, and the drone has already squeezed everything it can out of the near room. So it stays where the easy points are.

This is a data gap, not a skill ceiling. The proof is the specialist test. The same model, trained only on dumbbell maps, jumped from 10% to 95% success — and started confidently entering the doorway across all sorts of layouts. Nothing changed except what it was shown in training. The architecture and the reward were identical. So the model is perfectly capable of this behaviour; it just never learned it, because dumbbell maps made up roughly 0% of the original training mix.

How much should I mix in? I tested the dose — varying what fraction of training maps were dumbbells, then measuring success on a 1.4 m door:

share of dumbbell maps in training 0% 18% 45% 100%
success (1.4 m door) 10% 30% 50-60% 95%

The relationship is linear, with no threshold effect. There’s no cheap tipping point where a small dose suddenly unlocks the behaviour — to reach 95% you’d need almost 100% dumbbell maps, which would wreck performance on everything else. So “just mix in a fixed fraction” turns out to be the wrong tool for the job.

The decision

I’m recording this finding and moving on to SITL, which is the higher-value task right now. The reasoning:

  1. The curve is linear, so there’s no inexpensive fraction that buys most of the win.
  2. The dumbbell is an artificial stress test. Real doorways are ≥1.4 m, where 45% mixing already delivers 0.60 success without harming the rest.
  3. The maps we actually care about — multi-room layouts — are already solved (0.86 success).

If dumbbell maps ever become a serious requirement later, the right approach is a curriculum — temporarily show a lot of them, then taper them out — rather than a permanent fixed fraction.

4. Sources

  • This entry is the human-readable version. The full technical write-up lives in 28-two-chambers-distribution-gap.md.
  • Specialist-test and dose-sweep runs from the flight-quality sprint (parallel track, D2).

5. What’s next

  • Shift focus to SITL validation, which carries more value at this stage.
  • Keep the dumbbell maps on the shelf as a known stress test, to revisit with a curriculum approach if and when real-world layouts demand it.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR