What this is: A field note from the RL lab about checking whether our trained coverage policy keeps working when it is dropped into a two-stage “hybrid” flight: first a rule-based wall-follower traces the room edges, then the RL model takes over to explore the middle.
Why it’s here: It is a clean example of a question every robotics team eventually faces — will my learned controller still behave if the world it wakes up in is different from the world it trained in? The answer here is reassuring, and the reasoning is worth keeping.
Date: 2026-05-20
Glossary
Before the story, a few terms in plain English with everyday analogies.
- PPO (Proximal Policy Optimization) — the learning algorithm under the hood of our model. Think of it as a careful trainer: it improves the drone’s behaviour in small, conservative steps so it never lurches wildly away from what already works.
- Hybrid policy — a flight made of two different “brains” stitched together. One brain (simple rules) handles the first phase, then it hands the controls to a second brain (the RL model). The key worry with any hybrid is compatibility: the second brain has to be comfortable picking up wherever the first one left things.
- Wall-follower — a deterministic, non-learning routine that simply hugs the right-hand wall, like a person finding the exit of a dark maze by keeping one hand on the wall. It builds out the room’s perimeter without any RL involved.
- Handoff — the exact moment the wall-follower stops and the RL model takes over. The risky seam in a hybrid.
- visited grid — a 64×64 map where each cell is marked “already been here” or “not yet.” This is the model’s main sense of the world: it mostly looks at this map to decide where to go next.
- perimeter-init — an artificial way to pre-fill the visited grid so the cells next to walls are already marked “visited.” It imitates the state of the world right after a wall-follower has done its lap. We use it to fake the handoff during testing.
- coverage — the share of reachable floor cells the drone has visited. This is the mission’s main score.
- SWEEP-02 — our current production model.
- CAP3 — an experimental backup model, trained with one of its movement actions deliberately shortened (capped). More on it below.
- Gazebo — the 3D physics simulator (running on ROS2 Jazzy) where the sim team flies the drone in realistic rooms.
1. The setup: a drone that explores in two phases
Picture redecorating an unfamiliar apartment. The first thing you do is walk the perimeter — measure the walls, find the outlets and windows. Only then do you start working in the centre, arranging furniture and routing wiring.
The sim team built exactly this behaviour for the drone. Phase one: a simple wall-follower hugs the room’s edges (no RL — just “keep to the right wall”) and maps the boundary. Phase two: control passes to our RL model, which is free to explore the interior.
That two-phase plan is elegant, but it raises an honest question for the learned half. Our model was never trained this way. During training, the drone always started in a random cell with a completely empty visited grid — nothing seen, a blank slate. In the hybrid, the model instead wakes up to a visited grid that is already partly filled: a visible “donut” of visited cells running all the way around the walls.
So: will the model freeze up, or wander uselessly, when handed a half-finished map it has never encountered? That is the whole point of this note.
2. How I tested it
I wrote a compatibility check (scripts/compat_check_perimeter_init.py) that recreates the handoff without ever opening Gazebo:
- Take my normal training environment.
- Before each episode, manually mark as visited every cell that sits next to a wall (its 4-connected neighbours). This is the perimeter-init trick — it stands in for the lap the wall-follower would have flown.
- Place the drone in the centre of the room, not against a wall. (If it started by a wall it would effectively still be in the wall-follower’s phase; the centre is where the RL takes over.)
- Run the model for 1000 steps and measure how much additional floor it manages to explore.
- Repeat 33 times for each of 3 rooms — 99 episodes total.
- Then do the whole thing again for the backup model, CAP3.
The three rooms are the same SDF rooms the sim team builds in Gazebo:
- empty 6×6 — a plain 6.4 × 6.4 m room, nothing inside.
- pillar centre — the same room with a single column in the middle.
- two chambers — two rooms joined by a narrow corridor. The hardest of the three.
3. What I found
| Model | Empty room | With pillar | Two chambers | Average |
|---|---|---|---|---|
| SWEEP-02 | 75% | 75% | 66% | 72% ✅ |
| CAP3 | 61% | 63% | 55% | 60% ✅ |
| Threshold | — | — | — | > 40% |
Both models clear the 40% bar. But SWEEP-02 is the clear winner:
- It scores 12 points higher on average.
- It is far steadier — episode-to-episode spread of only 2–5%, versus 10–13% for CAP3.
- In its worst single episode, SWEEP-02 still covers 52%. CAP3’s worst drops to 16%, uncomfortably close to the 20% floor below which we’d have to retrain.
Why does the half-filled map not throw the model off? Because the model leans almost entirely on that 64×64 visited grid — a small convolutional filter reads the map and turns it into a sense of “where to go next.” A partly filled grid is just a natural extension of the patterns the model already saw in training; it is not an alien input. CAP3, by contrast, was trained with one of its movement actions artificially shortened, which made it less adventurous — it hangs back from the far corners and tends to stall.
The takeaway, stated plainly: SWEEP-02 stays in production, no retraining needed. CAP3 is parked in the archive in case real Gazebo flights later reveal that the long “fly forward” move truly must be limited.
We also decided not to add a wall mask to the observation. The sim team had offered to feed in an extra 64×64 “where the walls are” map. It turns out the model copes fine without it, so we keep the input simple.
4. What went into production
I extended the export gate, export/sweep02/validate_export.py. This is the green/red light that runs on every change to the export package (the model file and its metadata). It now has two modes:
default → checks the 5 standard evaluation maps
--perimeter-init → checks the hybrid handoff scenario
Both are green on the current SWEEP-02. If anyone ever swaps in a different model by accident, both tests will catch it and say “not this one.”
5. Where CAP3 came from
CAP3 has a backstory worth recording, because it explains why a second model exists at all.
The night before, the sim team tried running my model in Gazebo and got only 0.25 coverage where my mock-Gazebo had predicted 0.65. At first the suspect was the “fly forward until you hit a wall” action:
- In my environment, the drone crosses an entire corridor in a single action.
- In Gazebo, that same action runs as a long blocking loop and the drone only manages a cell or two before the safety guard says “stop.”
So I retrained a model that no longer relies on one action covering lots of ground at once — that became CAP3. But the next morning Aleks pinned down the real cause: it was a bug on the sim side, not a problem with my model. The plan pivoted to the hybrid approach, and CAP3 was kept only as a “might come in handy” spare.
6. A memory mishap, noted for next time
I kicked off three CAP3 retrains in parallel and the third got OOM-killed (the kernel terminated it for using too much memory). Three parallel trainings, each running 16 environments, add up to roughly 48 simultaneous simulations — more than the 16 GB session limit can hold. This is the same trap as on 2026-05-13. I reran the third one on its own and it finished in about five minutes. Lesson re-learned: count the total simulated worlds, not the number of training jobs.
7. Status
The Gazebo flight attempt is now unblocked; the sim team is waiting on the go-ahead from the coordinator. If the live Gazebo run goes well, the CAP3 branch retires for good. If it still comes back poor, we revisit — either promote CAP3, or train a fresh variant that resets directly from a perimeter-init state.