What this is: A dev-log entry about an “alignment sprint” — a focused effort to make our reinforcement-learning trainer agree with the realistic Gazebo physics simulator on one specific habit: how close the drone is allowed to fly toward a wall before it has to stop.
Why it’s here: It is a small-sounding rule with a large impact. When the trainer and the simulator disagree about something as basic as “how far from a wall is safe,” the trained drone freezes up and stops exploring. This entry walks through how we spotted the gap, how we proved the fix would work before spending any training time, and why we deliberately paused training to wait for one number.
Date: 2026-06-07 Ticket: rl-lab alignment sprint — wall margin
Glossary
A few terms show up a lot here. In plain English:
- 2D environment (env) — our lightweight “flight simulator on graph paper.” The drone learns to fly around a room drawn as a grid, where one cell equals 10 cm. It’s fast and cheap, perfect for training, but it’s a cartoon version of reality.
- Simulation (sim) — the neighboring system: the same drone, but inside Gazebo with realistic physics. Once a model is trained in the 2D env, it graduates to the sim. Think of the 2D env as a driving-school parking lot and the sim as the real road.
- OOD (out-of-distribution) — a situation the model never encountered during training, so it has no idea what to do and behaves foolishly. Like a student driver who only ever practiced in an empty lot, suddenly facing a roundabout.
- No-travel — the drone picks “fly forward” but doesn’t actually move anywhere. It’s stuck spinning its wheels.
- Margin (or wall margin) — how far from a wall the drone is required to stop. If the margin is 50 cm, the drone must halt with at least 50 cm of clearance ahead.
- Action mask — the list of “which moves are even allowed right now.” If flying into a wall isn’t permitted, the mask greys that option out before the drone can pick it.
- Raycast — how the drone “sees” walls: it shoots out sensor beams and measures where they hit, much like a lidar.
- Alignment — the whole point of this entry: getting two systems (the 2D trainer and the Gazebo sim) to agree on the same rules of the road, so a habit learned in one place still works in the other.
1. The problem: two simulators that disagreed
Our 2D trainer let the drone drive right up against a wall — as close as 10 cm. That felt fine inside the cartoon. But the realistic physics in Gazebo refuses to let the drone get that close. It applies a safety buffer and brings the drone to a halt 45–55 cm before the wall.
That difference created a real gap. The trained model wants to nose into that 45-cm band right next to the wall, because in training that was always allowed and even rewarded with progress. The real physics simply won’t let it. The result is the worst of both worlds: the drone keeps trying to move into space it can’t occupy, gets no movement in return, and ends up standing in place doing nothing — a classic no-travel freeze.
Here’s the mental picture. Imagine someone who learned to parallel park by bumping the curb every time. Put them in a car with automatic emergency braking that stops a half-meter short of the curb, and they keep stalling, confused about why the car won’t go where they expect. The driver isn’t broken and the car isn’t broken — they just learned different rules.
The goal of this sprint: teach the model from the start to keep about 50 cm of distance from every wall, so its habits line up with the real physics it will eventually face.
2. Checking before training (so we don’t train into a dead end)
Training runs cost time and compute. Before committing to one, I wanted evidence that the fix would actually help. So I took the current model — no retraining — and dropped it into the env with the new “stop 50 cm short” rule to see what happened.
Three things came out of that check:
-
With the old permissions, the model froze 95.77% of the time. That high a freeze rate confirmed the gap was not a minor edge case — it was real and large. Almost every step, the model wanted something the new rule forbade.
-
Just forbidding it from getting closer than 50 cm fixed the freezing entirely. By updating the action mask so the “fly into the wall band” option simply isn’t offered, the no-travel freezes dropped to 0%. And the quality of the coverage map jumped from 0.14 to 0.64 — without any training at all. That’s a strong signal the approach is sound. Training should clean up whatever is left, but the bulk of the win was already visible.
-
I worried the rule might trap the drone. My fear: a blanket “stay 50 cm from every wall” might seal off narrow doorways, leaving the drone unable to move between rooms. So I worked through the geometry carefully. The answer was reassuring — everything stays passable. The rule is about the wall directly ahead; the side walls of a corridor don’t block forward flight. Across all the test maps, 58–70% of the floor area remains reachable. And because the drone’s sensors still see walls clearly from 50 cm away, the coverage map still gets built completely — the drone “reads” the whole room without having to touch it.
That last point mattered a lot. It’s the difference between a safety rule that quietly cripples the robot and one that just makes it behave more like the real thing.
3. What I actually built
- Added a setting for how far to stop short of a wall (
wall_stop_cells). The important detail: the default is 0, which reproduces the old behavior exactly. Nothing existing changes unless you turn the new rule on. This is the boring-but-essential discipline of adding a feature without disturbing anyone who depends on the old one. - Verified nothing broke. All 19 existing tests plus 4 new ones pass — 23/23 green. The new tests cover the margin rule; the old ones prove the default still behaves identically.
- Wrote the rule’s exact geometry down in words and reviewed it with Aleks. This sounds like over-caution, but there’s history here: a similar small mismatch around a sensor once cost us about 1.5 months of confusion. Getting two systems to agree starts with both sides reading the same sentence and nodding. The wall-margin rule is conceptual — “stop a fixed distance before the wall ahead” — and the whole sprint hinges on the 2D env and the sim sharing the same understanding of that sentence.
4. What I’m waiting on (and why training is paused)
The simulation team is finishing its side and will send the exact margin number. It’s likely to be 60 cm rather than 50 — their safety buffer tends to run a bit larger. The plan is straightforward: once that number arrives and is confirmed, I launch the full training (four runs, for reliability), with the new margin baked in from step one.
The success criterion: freeze (no-travel) rate at or below 2%, down from the current 5–8%.
Training is intentionally frozen until then. This is a deliberate choice, not a delay. If I train against 50 cm and the real answer turns out to be 60 cm, I’ve just taught the model the wrong habit — and I’d be back to the same alignment gap, only now baked into a trained model. It is far cheaper to wait a day for one confirmed number than to retrain on a guess.
5. Takeaways
| What we learned | Why it matters |
|---|---|
| The wall-margin gap caused a 95.77% freeze rate | A “small” rule mismatch can almost completely stall the robot |
| Masking the forbidden moves lifted map quality from 0.14 to 0.64 with no training | The fix is fundamentally sound before any learning happens |
| 58–70% of floor area stays reachable; sensors still see walls from 50 cm | The safety rule doesn’t trap the drone or hurt mapping |
| 23/23 tests green, default behavior unchanged | The feature is additive and safe to merge |
The theme of this entry is alignment: a habit learned in the cheap, fast 2D trainer is only useful if it survives contact with the realistic Gazebo physics. Spotting where the two disagree — and proving the fix before spending training time — is most of the work. The training run itself is almost an afterthought once the rule is right and the number is confirmed.