What this is: A field note from training the ClaudeDrone indoor-mapping agent, where a safety improvement (keeping the drone away from walls) appeared to make map coverage worse — and how I figured out whether that drop was real.
Why it’s here: It’s a clean example of a recurring trap in robotics RL: a metric drops, everyone panics, and the first job is not to “fix” anything but to ask whether the metric is even measuring the right thing.
Date: 2026-06-08 Ticket: rl-lab indoor-coverage alignment / export
Glossary
- Alignment. Here it means teaching the drone to keep a steady, deliberate distance from walls instead of hugging them. Think of a new driver who learns to stop tailgating: safer, but they also stop slipping through tight gaps the way an aggressive driver would. “Aligned” behavior is calmer and more predictable — which is exactly what we want before handing the model to the simulation team.
- Coverage (mapped). The fraction of a room’s floor the drone has actually drawn onto its internal map. If a room has 100 reachable cells and the drone has charted 79 of them, coverage is 79%. It’s the headline “how much did you explore?” score.
- Model export. Packaging a trained model so another team can run it without retraining — like exporting a document to PDF so anyone can open it. Our export bundles three things: the model itself, a set of check tests, and a written “passport” describing what it does and doesn’t do.
- Fair export. An export whose accompanying numbers are honest — measured the right way, with known limitations spelled out in plain language, so the receiving team isn’t surprised later. The opposite of quietly shipping a flattering-but-misleading metric.
- PPO. Proximal Policy Optimization — the reinforcement-learning algorithm that trains the drone’s decision-making. Imagine teaching by lots of supervised trial-and-error, but with a built-in “don’t change your habits too abruptly between lessons” rule so learning stays stable. That stability rule matters later in this story.
1. What I wanted
The drone explores a room on its own, sweeping it with rangefinder beams — picture a blind person mapping a room by tapping a long cane in every direction and remembering where it hits.
In a recent change we taught it to stay about 0.6 m off the walls. Before that it would creep right up against them, and in the real simulator that caused it to “stick” — it would freeze near a wall, refusing to move, with travel commands going unexecuted. Keeping a margin solved the sticking. Good.
But the coverage number got worse: from 89% down to 79%. So Aleks lined up three questions, in order, and I worked them one at a time.
2. What I tried
Question 1 — is the number lying, or did it genuinely get worse?
My first suspicion was that we were being unfair to the drone. We compute “coverage” against all free cells, including a thin strip right up against the walls (that 0.6 m band) where the drone now refuses to drive. If it deliberately stays out of that strip, maybe we shouldn’t count those cells as “not covered” against it.
So I recomputed it more honestly: count only the cells the drone can physically see — even by a beam from a distance — while standing somewhere it can actually reach.
The number was not lying. It turns out the long-range rangefinder reaches 6.4 m, which spans the entire room. So the drone does see that wall-hugging strip — a beam fired from the center passes straight through it on its way to the wall. If it can see a cell, that cell legitimately counts as mappable. So 79% is real. The drop is real.
Question 2 — where exactly are we losing coverage?
Not everywhere. An empty room and a room with a pillar both scored great (83% and 91%). The losses concentrated in rooms with narrow doorways: a “two chambers through a gap” layout (62%) and multi-room layouts.
I checked the door width: 1.0 m. Maybe with a 0.6 m margin the drone simply can’t fit through? No — I computed reachability, and the far chamber is reachable. So this isn’t geometry. It’s behavior.
Aleks spotted the root cause and I confirmed it: the drone decides “can I drive forward?” based on the part of the map it has already drawn. Just past a doorway there’s still a blank patch — unexplored, undrawn. The drone treats that blank as “forbidden” and won’t step through the door. Back when it hugged the walls, it would push right up to the edge of the known area and “squeeze through” anyway. The new, calmer behavior makes it hesitate at the threshold.
Question 3 — fix it now, or ship as-is?
I decided to ship now. The reasoning: the whole point of this stage was to stop the drone from sticking at walls in the simulator — and that’s fixed. Narrow doorways are a separate, secondary problem for the next round. The simulation team is ready right now to verify the main metric (sticking) on our model, so making them wait for a doorway fix would block the thing that actually matters.
So I assembled an export pack for them: the best of the 4 trained models, a set of check tests for bit-for-bit agreement, and a model passport with honest numbers and a written description of the known limitation.
Bonus — a cheap test of the future fix
The obvious fix for doorways: let the drone decide “can I drive?” from the rangefinder reading (it can see through the door) rather than from the drawn map. I tested this cheaply, without retraining — I just swapped the rule on the already-trained models.
It got worse everywhere. That’s expected: each model learned under the old rule, so swapping the rule out from under it confuses it — like rearranging someone’s kitchen overnight and asking them to cook breakfast in the dark. So the fix is right in principle but needs retraining, not a hot-swap. I recorded that for the next stage, plus a note that we need to agree the rule in writing with the simulation team so both sides compute it identically.
3. What happened
- The coverage drop from 89% to 79% is real, not a measurement artifact — confirmed by recounting only genuinely visible cells.
- The loss is localized to narrow-doorway layouts, and the cause is behavioral (deciding from the drawn map), not physical (the drone fits through the door).
- The primary goal — no more wall-sticking — is achieved, so I shipped a fair export pack rather than holding it for a secondary fix.
- A no-retrain test of the proposed doorway fix degraded results, confirming the fix needs full retraining.
| Layout | Coverage |
|---|---|
| Empty room | 83% |
| Room with pillar | 91% |
| Two chambers through a gap | 62% |
| Overall (after alignment) | 79% |
| Overall (before alignment) | 89% |
4. Sources
- Coverage recomputation using line-of-sight visibility (only reachable, genuinely-seeable cells counted).
- Reachability analysis of the far chamber through the 1.0 m doorway.
- No-retrain rule-swap evaluation across the 4 trained models.
- Export pack: best model + bit-for-bit check tests + model passport.
5. What’s next
- Retrain with the rangefinder-based “can I drive?” rule, since a hot-swap on old models degrades behavior.
- Agree that rule in writing with the simulation team so the training environment and the real simulator compute it identically (bit-for-bit parity).
- Let the simulation team verify the main metric (no wall-sticking) on the exported model.