What this is: A field report from the rl-lab on an A/B experiment that swapped one kind of navigation hint for another — and found that the “more honest” version actually performed worse indoors.
Why it’s here: Negative results are some of the most valuable entries in any engineering journal. This one came with a clean diagnosis, so it tells us exactly what to try next instead of leaving us guessing.
Date: 2026-06-07 Ticket: TASK-RL-AM-3b
Glossary
- A/B test — running two training jobs that differ by exactly one idea, then comparing them fairly on the same maps and the same starting positions. Like baking two cakes that are identical except for one ingredient, so you can taste exactly what that ingredient does.
- BFS direction hint — BFS stands for Breadth-First Search, a classic way to find the shortest path through a maze by exploring outward one ring at a time. A “BFS direction hint” is a small piece of information handed to the drone that says, roughly, “the first step toward the nearest unexplored spot is this way.” Think of it as a compass that points along the real walkable route rather than straight through walls.
- Honest vs. lying compass — shorthand we use internally. The “lying” compass points straight at the goal even if a wall is in the way (fast to read, but misleading). The “honest” compass points along the actual route you’d have to walk (truthful, but it only knows about your next step).
- Coverage — how much of the floor plan the drone has actually mapped, expressed as a fraction. 0.90 means 90% of the explorable area has been seen.
- Instructive null — a negative result whose cause we fully understand. It’s worth more than a lucky win, because it permanently closes off a path we no longer need to walk down.
- Signal loudness — how easy it is for the neural network to tell apart the numbers coming in on its inputs. A faint, near-zero input is like a whisper in a noisy room: technically present, practically inaudible.
1. The one-paragraph version
We replaced a “lying compass” (which pointed at unmapped rooms straight through walls) with an “honest compass” (which points along the first step of the real route). On paper, honesty should win. Indoors, it lost: in apartment layouts the model now finishes mapping 80% of the floor in only 12 runs out of 50, down from 31. That is not a disaster. It is a paid lesson with a precise diagnosis, and it narrows the next move down to a single change.
2. The numbers, side by side
Animated runs for these scenarios live under /sim/rl/TASK-RL-AM-3b/.
| Scenario | v1 “lying” | v1.5 “honest” | Classic baseline |
|---|---|---|---|
| Empty room | 0.94 in 19 steps | 0.94 in 21 steps | 0.85 in ~260 steps |
| Room with a pillar | 0.88 | 0.90 | 0.85 |
| Two chambers joined by a doorway | 0.845 (8/10) | 0.62 (0/10) | 0.85 (10/10) |
| Apartments (5 layouts) | 31/50 reach goal | 12/50 | 50/50 |
Two things jump out. First, in the simplest spaces the honest compass holds its own — it even edges ahead in the room-with-a-pillar case. Second, the moment the layout gets genuinely tricky (two chambers connected by a single doorway), the honest version collapses to zero successful runs while the classic baseline sails through every time.
3. What actually happened — the diagnosis
We opened up one of the failed episodes and watched it step by step.
The good news is real: the old disease is gone. The v1 model used to twitch left-and-right between two unmapped rooms sitting behind walls, never committing to either. With the honest compass, that twitching genuinely disappeared. The cure worked.
The bad news is a brand-new symptom. Instead of twitching, the model now parks itself in a corner and spins in one direction for hundreds of steps — in the episode we inspected, 846 of them — going nowhere.
Here is why. The honest compass turned out to be too quiet. It only ever reports one of four possible first steps (north, south, west, east), and for a faraway target the strength of that hint is tiny — barely distinguishable from nothing, a whisper against the background noise of training. The old lying compass was wrong, but it spoke loudly and in every direction at once, so the model always had something to grab onto while it learned. We raised the truthfulness of the signal and, without meaning to, dropped its loudness. We lost more than we gained.
It is the engineering equivalent of replacing a bright-but-misleading road sign with an accurate one printed in faint grey ink. Accuracy is worthless if nobody can read it.
4. Why this is a manageable situation
A few reasons this null result is encouraging rather than alarming:
- The shared map vocabulary is untouched. The protocol we agreed on with the simulation team is fine — simulation keeps building its half exactly as before. The only thing we got wrong is how the hint’s loudness is encoded for the model to read. That is a small, contained surface.
- The next step is cheap and obvious. The natural follow-up is to reward the model for every step that shortens its route to the nearest unexplored area — the same honest routing logic, just delivered as a reward signal instead of as a faint compass reading. A nice property of this route is that it changes nothing in the data simulation sends us, so there is no re-negotiation to do.
- The v1 model is still here. It remains the strongest entry of the sprint, and the sprint’s headline goal was already met with it. Nothing was lost by trying the honest variant.
5. Where this got written down
Per Aleks’s standing instruction, the pivot and the sprint state were recorded in the team’s shared notes:
- An architecture decision record covering the shift from pure coverage toward active mapping, plus a new rule: agree on the map-parity protocol before training, not after.
- A living sprint page with the comparison table above.
- The rl-lab team card, with a “current sprint” section.
The takeaway for the journal: an honest signal still has to be a loud one. Truth that the network can’t hear is, for learning purposes, no truth at all — and that single sentence is worth more than the run-hours it cost us.