What this is: A “bench sprint” — we froze the drone’s brain (no retraining) and flew the exact same flight policy through seven different simulated rooms to see where it shines and where it stumbles.
Why it’s here: Before we move on to the next stage (closer-to-hardware SITL flying), we wanted an honest snapshot of what the current system can and can’t do. This is that snapshot.
Date: 2026-06-08 Ticket: “Stand” sprint (no retraining)
Glossary
- Policy / model — the trained “brain” that decides how the drone moves. In this sprint it stays frozen; we only test it, we never teach it anything new.
- Gazebo — the 3D physics simulator. Think of it as a flight-sim video game, except the player is our software, not a human.
- ROS2 (Jazzy) — the messaging system that lets all the drone’s software pieces talk to each other. Like a postal service for robot software.
- mavros — the translator between ROS2 and the flight controller’s native language (MAVLink). It turns “go forward” into the low-level commands an autopilot understands.
- SITL — Software In The Loop. Running the real autopilot firmware as a program on a PC instead of on a physical board. A dress rehearsal for real hardware.
- Out-of-domain (OOD) — a situation the drone never saw during training. Like a driver who only practiced in a parking lot suddenly being put on a highway.
- Coverage — how much of the room the drone managed to map. Higher is better.
- No-travel — a failure mode where the drone freezes and refuses to move. “No-travel 0%” means it always kept moving — a good thing.
- Perimeter sweep — after the main mapping mission, the drone traces the outline of the room, like walking the walls of a gym to measure it.
- Distance sensor — a small range-finder (“how far is the wall?”) feeding the autopilot. Real drones use these to avoid hitting things.
1. What we wanted
The goal was simple to state and tricky to pull off: take one already-trained flight policy and run it, unchanged, through seven very different rooms — from a plain empty box to a multi-room apartment — and write down exactly what happens.
No teaching, no tuning. Just observe. The point is to understand the policy’s real character before we invest in the next, more demanding stage.
Alongside the benchmark, the sprint bundled a few housekeeping tasks:
- Promote the v2 policy to be the default.
- Widen a doorway in one test world from 1.0 m to 1.4 m so the drone can actually fit through.
- Add four brand-new test rooms (an L-shaped corridor, a multi-pillar hall, a large open room, and an apartment).
- Add the perimeter-sweep phase that runs after the main mission.
- Wire up the distance sensors so the autopilot stops complaining about missing data.
2. What we did
Getting the drone off the ground
A few practical snags showed up before any benchmarking could start:
- The takeoff step wasn’t part of the default launch, so the bench script had to explicitly request an auto-takeoff. Without it the drone never arms and the system waits forever for a “ready to fly” signal that never comes.
- The autopilot refused to arm with a “PreArm: No Data” complaint, because the distance sensors were configured but nothing was actually feeding them numbers. We wrote a small forwarder node that takes the raw range-finder readings and republishes them in the format the autopilot expects. One subtle but important detail: when a sensor sees nothing (open space), we report the maximum range rather than skipping the reading entirely — skipping leaves the “No Data” error stuck on. After that fix the drone armed, took off, and held altitude correctly.
The perimeter sweep
The first version of the perimeter sweep used a fixed 8-second timeout per leg, which cut short the longer walls (around 5 m). We changed the timeout to scale with the length of each leg, and the traced rectangle came out clean and square.
A note on measuring coverage
The new rooms use wall masks generated by a fresh tool that draws walls more densely than the older test maps. Because of that, coverage numbers from the new rooms shouldn’t be compared one-to-one with the older empty/pillar rooms — it’s apples and oranges.
3. Results
Here’s the full run across all seven worlds. “MISSION” means the drone declared its mapping mission complete; “perimeter” shows how many of the four walls it traced.
| World | Mission | No-travel | Perimeter | Behavior |
|---|---|---|---|---|
| Empty 6×6 | Complete | 0% | 4/4 complete | Reference world; clean rectangular fly-around |
| Pillar (center) | Complete | 0% | 4/4 (center-return trimmed by timer) | Flew around the central pillar and mapped it |
| Two chambers | Not reached | 0% | — (no mission) | 1.4 m doorway is passable — drone flew from the left chamber into the right one. But it then looped inside the right chamber and never returned left, so coverage stayed low. The pass-through is a win; the incomplete tour is a limit of the policy on this unfamiliar layout |
| Multi-pillar | Complete | 0% | 4/4 (center trimmed) | Threaded past three pillars, mapped them, mission achieved |
| Large room | Complete (display) | 0% | 4/4 (center trimmed) | Flew the full 10×10 perimeter cleanly without getting lost, but didn’t fill in the interior. The “done” call fired early because the display map fills up too fast in a big room. Takeaway: the policy holds the boundary well, but interior coverage in large rooms is weak |
| L-corridor | Not reached | 0% | — (no mission) | Corner navigation works — both arms of the L were flown (horizontal arm, turn at the corner, then up the vertical arm). The corner was navigated successfully. Mission wasn’t declared (room is large and unfamiliar), but the stress-test goal was met |
| Apartment | Not reached | 0% | — (no mission) | The 1.4 m doorway between rooms A and B is passable — the drone moved from room A into room B, then got stuck in the lower part of B and never reached room C, so coverage was tiny. (First attempt hit a transient startup timeout; a clean retry came up in about 2 seconds) |
The big-picture finding
A few clear patterns came out of running one frozen policy across all seven rooms:
- No-travel was 0% in every single world. The drone never froze. That’s the single most reassuring result — the core safety behavior holds even far outside the conditions the policy trained on.
- Mission completion is reliable in familiar-style rooms (empty, single pillar, multi-pillar) but is either premature or out of reach in the big, unfamiliar rooms (large, L-corridor, apartment). In large rooms the internal “done” signal doesn’t match real-world coverage.
- Doorways of 1.4 m and wider are passable. Widening the doorway in task S1 paid off.
- Corner navigation works. The L-corridor proved the drone can round a corner and continue down the next arm.
- The perimeter sweep produces a clean rectangle (all four walls) in worlds where the main mission completes. The final center-return step sometimes gets trimmed by the runner’s 120-second window — cosmetic only, since the walls are already captured.
What the runs looked like
Each world produced a flight-track image and a mapping animation:
$DRONE_MEDIA_ROOT/sim/tracks/{track,map}_v2_<world>.{png,gif}
Because the benchmark ran headless (no on-screen window), there was no live video from it. So we recorded a separate set of GUI runs with the simulator window visible and captured them on screen. All seven worlds got a real-flight video — takeoff, then RL-driven mapping, plus the perimeter sweep where the mission completed:
$DRONE_MEDIA_ROOT/sim/tracks/video_v2_<world>.mp4
(One 130-second clip for the empty room and six 120-second clips for the rest, at 1200×996 resolution.)
4. Sources
- Engineering journal, simulation dev-log 29 (the distance-sensor forwarder fix).
- Flight tracks, maps, and videos:
$DRONE_MEDIA_ROOT/sim/tracks/.
5. What’s next
This bench was deliberately a “look, don’t touch” exercise, and it did its job: we now have an honest picture of where the policy is strong (safety, boundary-following, familiar rooms, doorways, corners) and where it’s weak (interior coverage in large rooms, completing tours in unfamiliar multi-room layouts).
The next stage moves toward SITL — running the real autopilot firmware in the loop — to get closer to how the drone will behave on actual hardware. The weak spots found here give us a clear list of what to watch for there.