What this is: A development log from the simulation team describing a two-stage flight strategy: a hand-written “wall follower” traces the room’s perimeter, and once the loop closes, a learned policy takes over to explore the middle.
Why it’s here: It documents a real pivot in how our drone explores an enclosed space — away from mowing the room in stripes, toward first pinning down the walls and only then filling the interior. It is a good window into how deterministic logic and machine learning can share one mission.
Date: 2026-05-20 Ticket: TASK-062 (Wall Follower + RL Hybrid)
Glossary
A few terms up front so the rest reads smoothly:
- Wall follower — a simple, rule-based behaviour where the drone always keeps a wall on one side of itself (we use the right-hand rule, like walking through a maze with your right hand always touching the wall). Follow that one rule long enough and you will eventually trace the entire outer boundary of a room.
- RL (reinforcement learning) policy — a behaviour that was learned by trial and error in training rather than coded by hand. Think of it as a driver who got good through thousands of practice laps instead of being handed a rulebook. We train ours with PPO, a standard learning method.
- Hybrid — using both at once, but at different times. The drone is deterministic and predictable while hugging the walls, then switches to the learned policy for the harder, more open part of the job.
- Fallback policy / mode switch — the ability to fall back to a single, simpler behaviour. If something misbehaves, we can run “walls only” or “learned policy only” instead of the full hybrid, which makes debugging far easier.
- Occupancy mask / wall map — a grid of cells, each one either “wall here” or “not seen yet,” like graph paper where we shade in squares as the drone discovers them.
- Perimeter complete — our way of deciding the drone has gone all the way around the room once and is back where it started.
1. What we wanted
The trigger for this work was a clear directive from Aleks: the drone should run up against the walls and lock them down first. Until then our exploration plan looked like a lawnmower — sweep the room in parallel stripes. That is fine on paper, but it treats the walls as just another part of the floor, and the walls are exactly the part we most want to nail down early.
So the goal became a two-act plan:
- Act one — find and trace the walls. Fly to the nearest wall, keep it on the right side, and follow it all the way around the room, recording every wall surface we touch.
- Act two — fill the middle. Once the perimeter is mapped and we are back near the start, hand the controls to the learned RL policy to cover the open interior.
I also wanted to keep my own sanity during testing, so the plan had to be switchable: full hybrid by default, but with the option to run just the wall follower or just the learned policy on demand.
2. What we did
I split the work into five pieces, each small and independently testable.
2.1 The wall follower itself
This is a little state machine — at any moment the drone is in exactly one of a handful of “modes,” and it moves between them based on what its distance sensors report:
- Find a wall — every sensor reads far away, so just fly forward until we bump into something.
- Approach a wall — the front sensor drops below the target wall distance (about 0.6 m), so turn 90° to the right and start following.
- Follow the wall — keep the right-hand wall at the target distance. If the front gets too close, we have hit an inner corner, so turn left 90°. If the right-hand wall suddenly falls away, we have reached an outer corner, so steer back toward it.
- Corner turn — execute the 90° rotation cleanly, then resume following.
One subtle but important guard: we refuse to declare the perimeter “done” for the first several seconds of flight. Without that guard, the drone would think it had completed the loop the instant it took off (it is still sitting at the start point, after all). A short minimum-flight timer fixes this neatly.
The wall follower never outputs raw velocities — it outputs target positions and headings. That is a deliberate carry-over from an earlier refactor: the bridge publishes a position setpoint to mavros, and the drone’s flight controller handles the rest. Crucially, the heading actually changes, which earlier work had confirmed end to end.
2.2 Building the wall map
As the drone follows the perimeter, a second component records what it sees onto a 64×64 grid covering the room (roughly 6.4 m on a side, with each cell about 10 cm). For each of the six distance sensors (spread evenly around the drone), it:
- ignores readings that are saturated (too far to be a real wall),
- works out where in the room the detected wall surface actually sits, given the drone’s current position and heading,
- and shades in the matching grid cell.
The result is a slowly filling-in picture of the room’s outline.
2.3 The phase controller
This is the conductor. On every tick it looks at the sensors and the drone’s pose and decides which act we are in:
- Wall-follow phase — update the wall map and the “places I’ve been” grid, check whether the perimeter is complete, and ask the wall follower for the next move.
- RL-explore phase — hand off to the existing learned-policy pipeline that turns observations into actions.
When the switch happens, it logs a clear marker so we can see exactly when and why control changed hands, including how many wall cells were found and how many laps were flown.
2.4 Wiring it into the bridge
The bridge node is where everything meets the rest of the ROS2 (Jazzy) system. I pulled in the three new components and added launch parameters so the behaviour is configurable at startup without touching code. The most important one is the mode parameter:
| Mode | What it does |
|---|---|
hybrid (default) |
Full plan: trace the walls, then switch to the learned policy |
wall_follow_only |
Just the perimeter trace — handy for debugging the wall logic alone |
rl_only |
Just the learned policy, the way things behaved before this ticket |
Having rl_only as an escape hatch means we never lose the old behaviour while iterating on the new one.
2.5 Tests
Finally, a suite of twelve fast unit tests, all using mocked sensor data so they run in a fraction of a second without any simulator. They cover the interesting transitions:
| Test | What it checks |
|---|---|
| Finds a wall | All sensors far → fly forward, heading unchanged |
| Turns right at a wall | Front sensor close → turn right and start following |
| Follows the right wall | Right sensor at target distance → glide along it |
| Inner-corner turn | Front gets too close → turn left |
| Perimeter complete | Back at start after enough flight time → done |
| Perimeter not complete (too early) | Not enough flight time yet → not done |
| Map records a hit | A close reading → the right grid cell gets shaded |
| Map skips out-of-range | A far reading → nothing recorded |
| Map across multiple sensors | Two sensors hitting → at least two cells shaded |
| Starts in wall-follow phase | The conductor begins in the right act |
| Switches to RL on perimeter complete | Loop closed → control handed over |
| Updates “visited” during wall phase | Coverage bookkeeping keeps running |
3. Results
All twelve tests passed, and they run in well under a tenth of a second because they need no simulator at all — pure Python with faked inputs. That speed matters: it means I can re-run the whole suite after every small change without breaking my stride.
What this proves is that the logic — the state transitions, the map-building math, the phase switch — is sound in isolation. What it does not yet prove is real-time behaviour in Gazebo, which is the next frontier (more on that below).
The time budget tells its own story. I had penciled in roughly four hours across the five pieces and finished in about an hour and twenty minutes:
| Sub-task | Budgeted | Actual |
|---|---|---|
| Wall follower | 1.5–2 h | ~25 min |
| Wall map builder | 45 min | ~10 min |
| Phase controller | 30 min | ~10 min |
| Bridge integration | 30 min | ~30 min |
| Mock tests | 30 min | ~5 min |
| Total | ~4 h | ~1 h 20 min |
The big overestimate was on the tests — mocked unit tests are genuinely cheap to write. And the first version of the wall follower is deliberately a simple state machine; the fiddly corner cases are reserved for tuning once we see the drone fly for real.
4. Sources
- The directive that started the pivot: Aleks’s call to make the drone run up against the walls and fix them in the map first.
- An earlier control refactor that established position-setpoint output and confirmed that heading changes propagate through to mavros — the wall follower builds directly on it.
- The existing learned-policy pipeline (PPO-trained), which the hybrid hands off to for the interior phase.
5. What’s next
The code is ready; the real test is live flight. A few things I am specifically watching for:
- A live Gazebo / SITL smoke test. Earlier learned-policy-only runs were difficult to get stable, and the hybrid has not yet flown in real time. The things I expect to wrestle with are the wall follower wobbling at corners (rapidly flipping between “follow” and “turn”), the perimeter-complete check firing slightly too early or too late, and making sure the conductor is genuinely driving the main loop.
- A fresh learned policy from the RL team. The interior phase currently loads an older trained model; a newer one is needed for a fair test of the RL act.
- A cleaner wall map. Right now the map is raw points with no smoothing, so the learned policy might see gaps where the perimeter is actually solid. Tidying the map is a candidate improvement once the basics work.
If the perimeter never closes within a reasonable window, the fix is in the wall follower’s turn angles and corner detection. If the handoff happens but the learned policy drifts in the open middle, the answer is either a retrained policy or feeding the wall map into the policy’s observations.