What this is: A post-mortem of the very first time we connected our trained navigation policy to a real simulated drone — and watched it fly straight into a wall. Why it’s here: Failures teach more than successes. This walks through five separate bugs that conspired to crash the drone, and how each was fixed. It’s a textbook example of why integration testing exists.
Date: 2026-05-19 (flight) → 2026-05-20 (analysis) Ticket: TASK-059, attempt #1
Glossary
- Policy — the trained “brain” that decides what the drone should do next (turn, go forward, hover). It looks at sensor readings and outputs an action number.
- Bridge — the piece of software that connects the policy to the actual drone. It feeds sensor data into the policy and turns the policy’s decisions into motor commands.
- RCA (root-cause analysis) — a structured hunt for the real underlying reasons something broke, not just the surface symptom. Think of a doctor diagnosing the disease behind the cough, not just prescribing cough syrup.
- QoS (Quality of Service) — in ROS2, a set of “delivery rules” for messages between programs. Two programs must agree on the rules, or messages get silently dropped — like a phone call where one side speaks French and the other only listens in English.
- odom (odometry) — the drone’s running estimate of where it is and how it’s moving (position, velocity).
- SITL (Software In The Loop) — the real flight-control software running on your computer instead of on a physical board, so you can test flight behavior without a real drone.
- mavros — the translator between ROS2 and the flight controller. It carries commands down and telemetry back up.
- Stale observation — sensor data that is old and no longer reflects reality. The drone “thinks” it’s somewhere it isn’t.
- VL53L0X — a small distance sensor (laser time-of-flight). In our setup it measures how far walls are.
1. TL;DR — what happened
During the first integration test, we connected the navigation policy to a real Gazebo world with SITL flight control. The drone took off cleanly and hovered at 2 metres. Then everything went wrong:
- The policy started receiving stale position data (a QoS rules mismatch meant odometry messages were silently dropped).
- The policy also received corrupted distance readings (a sensor bug turned “open space far away” into “wall right next to me”).
- Believing it was in an empty, unexplored room with clear space ahead, the policy chose to fly forward.
- The wall-stop safety threshold was set so tight (15 cm) that the drone couldn’t brake in time.
- The drone flew across the room, hit the wall, and climbed vertically to 6.21 m before we manually shut it down.
We found five distinct root causes, applied four code fixes, and flagged the remaining work for a joint validation run.
2. Timeline
| Real time | Event |
|---|---|
| 23:25:44 | Takeoff command issued |
| 23:25:57 | Hover reached — z = 2 m, mode GUIDED, armed |
| 23:27:53 | Bridge launched |
| 23:27:57 | Bridge ready — warning logged: odom QoS incompatible |
| 23:28:18 | First “fly forward until wall” safety cap triggered (first flyaway) |
| 23:28:30 | Policy stuck in a rotate-in-place loop, coverage flat at 0.000 |
| 23:29 | Operator alert: “drone in the corner at z = 6.21 m” |
| 23:29:06 | First fix applied (odom QoS) and bridge restarted |
| 23:30:13 | Second bridge run — no warning, but coverage still stuck at 0.000 |
| 23:31 | LAND command sent; drone disarmed |
| 23:32 | Final state: disarmed, reported pose stuck, real altitude 6.21 m |
| ~23:45 | Fifth root cause identified (the sensor corruption bug) |
3. The five root causes
Root cause #1 — the launch script spawned Gazebo twice
Our “full launch” mode starts several components, each in its own terminal pane. One pane was supposed to start Gazebo. But another pane, which launches the ROS2 stack, also started a Gazebo instance of its own. Two simulators ran in the same partition at once, which confused the SITL communications and made time synchronisation jump around.
Fix: The launch file now accepts a launch_gz flag. When Gazebo is already being started elsewhere, the launch script automatically passes launch_gz:=false so the second instance is never spawned. It also exports the correct world name so the right room loads either way.
Root cause #2 — the launch file never started takeoff
We had assumed the main launch file would automatically take the drone off the ground. It doesn’t — it starts the sensor and control nodes, but not the takeoff routine.
Workaround: Run the takeoff command separately. The event-driven takeoff node handles the whole sequence: connect to the flight controller, switch to GUIDED mode, arm, climb to 2 m, then hover. This works reliably.
Decision (deferred): Whether to fold takeoff into the main launch file is left for a later ticket — some tests deliberately don’t want auto-takeoff.
Root cause #3 — QoS mismatch on odometry (the main cause of the flyaway)
This was the big one. mavros publishes odometry using “best-effort” delivery rules. The bridge subscribed expecting “reliable” delivery. In ROS2, this exact combination is incompatible — the publisher and subscriber simply can’t talk, and messages are silently dropped.
The result: the bridge received zero position updates. The policy’s internal sense of position stayed frozen at the origin (0, 0), so its map of where it had been collapsed to a single cell. Coverage read as essentially zero.
To the policy, this looked like: “I’m in an empty room, everything is unexplored, and there’s open space straight ahead.” So it decided to fly forward to explore.
Fix: Subscribe to the odometry topic with the sensor-data QoS profile (best-effort), matching what mavros publishes. We also added a property to track whether any odometry has actually been received.
Lesson for future work: every
/mavros/*sensor topic uses best-effort delivery. A subscriber that defaults to reliable will fail silently. Always use the sensor-data QoS profile for these topics. We confirmed this against the official ROS2 documentation: “for the reliability QoS policy, all combinations are compatible except when the publisher uses Best Effort and the subscriber uses Reliable.”
Root cause #4 — the wall-stop threshold was far too tight
The “fly forward until collision” behaviour kept moving the drone until a front sensor reported a wall closer than 15 cm. Two problems compounded here.
First, the front distance sensor maxes out at 2 m, and our processing clamped that further — so the bridge always believed any obstacle was at least 1.2 m away. With a 15 cm stop threshold, the drone would only try to stop once it was 15 centimetres from the wall.
Second, at 0.3 m/s the drone covers 15 cm in half a second. By the time it “decided” to stop, it was already in the wall. Over the full safety window, that forward motion adds up to roughly the entire width of the room.
Fix: Raise the wall-stop threshold from 0.15 m to 0.50 m. At 0.3 m/s, half a metre gives about 1.67 seconds to brake — enough margin to actually stop.
We also confirmed a flight-controller safety net we weren’t using: in GUIDED mode, the vehicle stops on its own after 3 seconds with no command. But the bridge kept streaming forward commands continuously, so that safety never kicked in.
Root cause #5 — the sensor monitor turned “far away” into “right here”
The sensor monitor had a subtle bug. When a distance sensor read “infinity” (meaning: open space, nothing within range), the code simply skipped the update instead of recording it. Uninitialised channels then defaulted to a reading of 0.0 — which the bridge interprets as “obstacle at zero distance.”
So any direction looking out into open space reported a wall pressed right against the drone.
Live evidence at the moment of the crash: the rear-facing channels read 0.0 while their raw sensors reported infinity. The drone was sitting in a corner with open space behind it — but the policy “saw” walls there. Even after the odometry fix, the policy still received these corrupted distances. Seeing “obstacle at 0 m” behind it, it refused to back out and spammed a rotate-in-place action for 200 steps.
Fix: Initialise all channels to the sensor’s maximum range instead of zero. In the callback, clamp any infinite or over-range reading to the maximum range rather than dropping it, and always publish. The same fix was applied to the altitude sensor.
Side benefit: the sensor now publishes at a steady, honest rate, which makes the data far easier to observe and debug.
4. Patches applied
| Component | Change |
|---|---|
| Main launch file | Added launch_gz flag wrapping the Gazebo start |
| Launch script | Exports the world name, auto-passes launch_gz |
| Observation builder | Sensor-data QoS for odometry; tracks first odom received |
| Sensor monitor | Clamps infinity to max range, initialises to max, always publishes |
| Bridge node | Wall threshold 0.15 → 0.50 m; added odometry-age and safe-box pre-flight checks |
| Failure handling | Added stale-odom and out-of-bounds triggers, plus recovery logic |
5. Safety net added to the bridge
Before the policy decides anything, the bridge now runs three pre-flight checks:
- If the position data is older than 1 second, hover and skip the decision (this skip does not count against the step budget).
- If the drone has drifted outside a safe box (half the room size plus a small margin), hold a permanent hover until a human intervenes.
- If the drone is hovering but conditions have recovered (fresh position data, back inside the box), exit hover and resume.
6. Validation still pending
The next step is a joint mock test run to confirm the patches don’t break anything that already worked:
- Pull the patched bridge code.
- Run a sweep of several rooms and seeds in the mock environment.
- Compare coverage against the earlier synthetic baselines (0.951 / 0.950 / 0.937).
- If they match, the patches are cleared for a real Gazebo retry.
The mock environment passes data directly in Python, so it never reproduced the QoS or sensor-corruption bugs in the first place — which is exactly why they slipped through. The patches only add safety checks for non-mock failure modes, so the mock results should hold steady.
7. Open questions
- When position data goes stale for more than a second, what should the bridge do? Current proposal: publish a zero command and skip the decision, without counting it against the step budget.
- The 0.15 m wall threshold came from training in the mock environment, but real Gazebo needs at least 0.5 m. Is that a policy specification or a bridge config value? If the policy genuinely expects 0.15 m, we have a training-versus-reality gap to close.
- Did the policy’s training data ever include corrupted distance readings (0.0 in unobstructed directions)? If not, the bridge must guarantee valid observations before every decision — which is exactly what the sensor-monitor fix does.
8. Lessons learned
- Integration testing catches what synthetic testing can’t. Our mock test passed with a perfect bit-exact match, yet it never touched real QoS settings, real infinity readings from the sensors, or the launch-file side effects. Those only appeared with real Gazebo and real mavros.
- Sensor pipeline integrity matters before any decision. Garbage in, garbage out. The bridge should validate observations and warn on suspiciously tiny readings.
- A real-world safety net is mandatory. The mock never simulates stale data or out-of-bounds drift. The bridge must have pre-flight checks always.
- Two-stage launch beats all-in-one. Starting the simulator separately from the ROS2 stack avoids surprise duplicate spawns.
9. Next actions for attempt #2
- Patches applied (done).
- Joint mock test (waiting on the validation run).
- Operator confirmation that it’s safe to reset.
- Clean restart with patches in place.
- Bring up Gazebo, SITL, and mavros in the empty 6×6 room.
- Run takeoff to hover.
- Launch the bridge with patched code.
- Watch the action distribution, coverage growth, and confirm no flyaway.
- Record a full episode.
- Write up the results.