What this is: An engineering-journal entry from a long simulation session — first an overnight “watch” over a sequence of training/eval runs (labelled A through F), then a follow-up day of runs that finally answered one blunt question: is the drone merely not crashing, or is the model genuinely flying and covering the room?
Why it’s here: It’s a candid look at how we separate “the aircraft survives” from “the policy is good,” how we hunt down control-stack fights between two systems both trying to steer the drone, and how a sneaky autopilot version mismatch silently swallowed half our tuning for weeks.
Date: 2026-06-06 22:00 → 2026-06-07 (day) Ticket: v2 night watch / batch runs A–F
Glossary
Before the dense parts, here’s the plain-English version of the jargon. If you already live in this stack, skip ahead.
- Night-watch run. Exactly what it sounds like: I sit with the simulator running overnight, checking in every 20–30 minutes, reading logs, and fixing problems the moment they appear instead of waking up to a dead run and no idea why. Think of it like a night nurse doing rounds — the patient (the simulated drone) is stable, but you keep checking the monitors.
- Batch runs A–F. Each “run” is one full attempt of the drone flying in the simulated room under a specific configuration. We label them A, B, C, … so we can talk about them precisely. Each letter is “same experiment, one knob changed” — like baking the same cake six times and only adjusting the oven temperature each time so you know exactly what caused the difference.
- SITL (Software In The Loop). The real flight-control software (the autopilot brain) runs as a program on the computer instead of on a physical flight board. The drone is fake, but the autopilot logic is the real thing. It’s a flight simulator where the pilot’s actual brain is wired in, just without a body.
- Gazebo. The 3D physics simulator. It models the room, the walls, gravity, and the drone’s sensors so the autopilot thinks it’s really flying. It’s the “world” that wraps around the SITL “brain.”
- mavros. The translator that lets ROS2 (our robotics middleware) talk to the autopilot using its native language (MAVLink). Picture a bilingual interpreter sitting between two people who’d otherwise just stare at each other.
- ArduPilot. The open-source autopilot software — the actual flight-control brain running under SITL.
- coverage@338. A score for how much of the room the drone has “painted” (visited/mapped) by a fixed early checkpoint (step 338). It’s our quick early read on whether the policy is doing useful work, the way you’d glance at a runner’s pace at the first kilometre instead of waiting for the finish line.
- escapes/step and safety guard. “Escapes” is how often the drone wriggles in ways that look like it’s stuck and thrashing; the “safety guard” is the protective layer that pushes the drone back when it gets too close to a wall. Lower numbers are better for both.
- PPO. The reinforcement-learning algorithm that trained the flying policy. You don’t need its internals here — just that it’s the method that produced “the model.”
1. What we wanted
A reviewer (plus Aleks) drew a sharp line that the whole session orbits around: “the drone doesn’t crash” is not the same as “the model flies.” A drone can survive a run by doing nothing useful — hovering in a corner, parking against a wall — and still pass a naive “no crash” check. So the review handed me three concrete acceptance criteria to judge an actual flight by:
- Coverage@338 ≥ 0.25 — the model must paint at least a quarter of the room by the early checkpoint.
- Escapes/step ≤ 3% — it shouldn’t be thrashing.
- Safety-guard time ≤ 5 seconds per minute — the protective layer shouldn’t be doing the flying for it.
For reference, the same model in its native training environment (identical room) scores 0.351@338 and reaches 0.95 coverage by the long horizon. That’s the bar: the live SITL+Gazebo stack should reproduce something close to that, not just keep the aircraft intact.
The directive for the night was simple and old-school: watch in 20–30 minute rounds, keep logs, fix issues on the spot, and research solutions online as problems came up.
2. What we did
2.1 The overnight sequence (runs A→E)
I worked through configurations one change at a time. Here’s the night’s chronology, verdicts read at the @338 checkpoint:
| run | config | fate | verdict @338 |
|---|---|---|---|
| old | pre-v2 | crash within 5 min | 0.113 / 6.8% / 34s — ✗✗✗ |
| A | parity fixes | survived to 338 ✓ | 0.255 ✓ / 5.0% ✗ / 20s ✗ |
| B | + AND-gate stuck detector, yaw 90°/s | crash, AngErr=61 @12 min | tug-of-war v2: guard Twist @50Hz vs executor position @10Hz |
| C | + arbitration on /safety/active + retreat | stall @25 min | hysteresis dead-band [0.8, 0.9] — my bug |
| D | + retreat-until-release | alive, no crashes | 0.140 ✗ / 3.6% / 7.8s — capped by a 0.8m ring (44% of the room) |
| E | + margin-set 0.5/0.6/0.75 | whole night, no crashes | 0.155 ✗ / 2.4% ✓ / 2.6s ✓ |
Run A was the first encouraging sign — it cleared the coverage bar (0.255) but failed the other two. Then B and C each taught me something the hard way (more below). E was the keeper: it ran the entire night without a single crash and, more importantly, satisfied two of the three criteria (escapes and guard time both green).
Run E’s trajectory over time: 0.155@338 → 0.272@888 → 0.383@2000 → 0.443@2400. The growth was monotonic over roughly four hours, and my interventions tapered off as it went (16 → 9). So the model is learning to cover the room as the episode unfolds; it just starts slow at the early checkpoint.
2.2 The control-stack fight (run B)
Run B crashed with a large attitude error, and the root cause was a fight I’ll call the control-stack tug-of-war. Two systems were both streaming commands to ArduPilot at once: a “position maintenance” bridge and a “safety push-back” guard, at different rates (10Hz vs 50Hz). ArduPilot got yanked between modes and lost attitude control until it tumbled.
Mental image: two people grabbing the same steering wheel, one nudging gently to hold a lane, the other yanking hard to dodge a wall — the car swerves and spins because nobody is fully in charge.
The fix was arbitration: a latched /safety/active flag. When the guard takes over, the maintenance bridge freezes (stops fighting), the guard slowly retreats the drone away from the nearest hazard channel, and on release the bridge re-initialises its target to the drone’s current pose so there’s no jerk. One hand on the wheel at a time.
2.3 The hysteresis bug (run C)
Run C stalled, and that one was my own mistake. The retreat behaviour was only active during an active alarm, not until full release — so the drone parked forever inside the hysteresis band, never quite triggering and never quite clearing. The rule learned: retreat must run until release, not just while the alarm is asserted. Otherwise you get an eternal parking spot in the no-man’s-land between thresholds.
3. Results & key findings
3.1 Night findings
- Control-stack tug-of-war (run B). You cannot simultaneously stream position-maintenance commands and safety-Twist commands — ArduPilot oscillates between modes until an attitude crash. Resolved with the latched-arbitration scheme above.
- Hysteresis dead-band (run C). Retreat must run until full release. Anything less parks the drone in the hysteresis band indefinitely.
- A dropped
infreading froze the pipeline. The observation builder was discardinginfrange readings — butinfis a legitimate reading here: the room’s diagonal (8.7m) exceeds the sensor’s range (6.4m), so “I see nothing within range” is the correct answer, not an error. Dropping it froze the data timestamp, and a freshness gate then blocked the model from predicting forever. Fix: convertinfto a capped value and keep the timestamp fresh. - A geometric ceiling on coverage. The coverage metric is computed from a free-space mask plus a distance transform. The safety ring we’d configured (a 0.95m keep-out band) mathematically capped achievable coverage: the wider the band, the lower the maximum possible score. At one setting a 0.8m floor locked 44% of the room away from a policy that was specifically trained to paint along the walls. That’s a structural explanation for the failing coverage@338 — the geometry, not the brain, was the limit at that setting.
- The stuck detector should watch the world, not the action. Research on active-mapping agents pointed the same way: detect “stuck” from the environment state (how far the drone moved, how much it yawed, how coverage changed over a window), never from the action the policy chose. Watching the action punishes a policy that legitimately rotates in place ~83% of the time. Confirmed empirically: escapes dropped 6.8% → 2.4%.
- A simulator sensor glitch self-healed. Around 05:00, the Gazebo GPU-lidar channels dropped to 0 Hz for about 3.5 hours, then recovered on their own to 4.7 Hz. Suspected EGL graphics degradation after 8 simulator restarts over the night. Recommendation: reboot the machine before long sessions. A side note for the backlog: the sensor monitor’s 10Hz timer can mask a dead sensor by re-publishing cached data — it needs to fail loud instead.
3.2 Morning decisions → run F
After the night, Aleks set the direction for run F:
- The 0.4m clearance is sim-only. Wall aerodynamic effects aren’t modelled in Gazebo, so for real hardware this gets recomputed from propeller diameter plus the diagonal frame. We don’t go below 0.4m — near-wall states there don’t transfer reliably to the real aircraft. New set: 0.4 / 0.5 / 0.65.
- Rotations via a raw setpoint mask (zero velocity + yaw rate only), open-loop duration derived from angle ÷ rate, observations sampled on a timer (not on an arrival event), with achieved-vs-commanded calibration logged every 25 rotations.
- Velocity-gated arrival: count “arrived” only when position error is small and the drone is nearly stationary; tighten the settle threshold over time; log drift telemetry every 50 transactions so we can lower the threshold if a systematic drift appears.
3.3 The day sequence (runs F→H) and the autopilot surprise
The day picked up from run F and ran the sequence to H. Day chronology, verdict columns = coverage / escapes / guard-seconds-per-minute @338:
| run | config | verdict @338 | fate |
|---|---|---|---|
| F-1 | floor 0.4 == margin 0.4 | — | 11 safety triggers in 18 steps on an oblique near-wall channel (0.395–0.400m), guard tug-of-war at the boundary → restart per checklist |
| F-2 | floor 0.45, coherent set | 0.178 ✗ / 2.1% ✓ / 4.2 ✓ | every move hit an arrival timeout (15s), stopping 0.3–0.75m short |
| F-3 | + Z-capture at release | — | Z error 0.18→0.02m ✓, but horizontal median 0.065 unchanged: Z-coupling theory disproven; 9/100 timeouts |
| F-4 | + carrot streaming | — | carrot alone didn’t help (median 0.051) → parameter dig → 🔴 autopilot version finding; with a live guidance-option fix: median 0.102, p90 0.156, 0/110 timeouts |
| G | + indoor parameter-file audit | 0.161 ✗ / 1.8% ✓ / 10.1 ✗ | timeouts gone ✓, but 135/135 triggers parked on the line (0.437–0.450): margin==floor again — now the drone reaches the wall and parks exactly on the trigger line |
| H | floor 0.40, margin 0.45 | 0.139 ✗ / 1.5% ✓ / 1.4 ✓ | 12/12 triggers honest (0.395–0.400), guard topic closed |
The headline finding of the day was the 🔴 one: the ArduPilot tree we’d been flying was a development branch (4.8.0-dev), not the Copter 4.5 release we assumed. Several parameters had been renamed between versions, and — crucially — unknown parameter names in the config file are ignored silently, with zero warnings. That meant a whole floor of our indoor tuning file was inert. It had been quietly doing nothing, including a fix from a much earlier attempt that was supposed to cure a chronic vertical lag (~0.18m). For weeks we’d been “fixing” things the autopilot never read. Aleks’s call: stay on the dev branch (it matches the night’s baselines and is closer to the future hardware) and do a full parameter audit.
Two more day findings worth keeping:
- One guidance option was poison on this autopilot branch when streaming. A particular guidance flag, combined with 10Hz command republishing, forced the autopilot to restart its motion-smoothing curve every 100ms — so the speed ramp never finished, and the drone crawled horizontally at ~0.05 m/s regardless of the commanded speed. Switching that flag off (the native position-control path for offboard control) unblocked it instantly in a live A/B test.
- Commanded speeds were being legacy-ignored. The mode speeds (0.05 / 0.15 / 0.30) hadn’t reached the autopilot for ages. “Carrot streaming” — a movable target placed a clamped 0.3m ahead of the drone’s pose, advanced by time × speed — finally made them real: horizontal p90 of 0.156 now matches the commanded 0.15.
- margin ≠ floor is a rule, learned twice. When the safety margin equals the floor, the executor parks the drone exactly on the guard’s trigger line and the guard fires endlessly. A 0.05m buffer is mandatory: floor 0.40 / margin 0.45 / gate 0.55 / wall-target 0.70.
And the honest punchline (finding 5 of the day): the cleaner the mechanics got, the more clearly we could see the model’s ceiling. Coverage@338 across the clean runs went 0.178 → 0.161 → 0.139 with near-perfect execution. The earlier timeouts had been accidentally “sweeping” the room as the drone drifted around; once execution got precise, the inefficiency of the policy’s sweep pattern was exposed. The limit at this point is the model, not the plumbing — which is exactly the question we set out to answer.
4. Tools built (all in the help-scripts folder)
- chain_monitor.py — a watchdog over the input/output command chains. It caught all three night incidents before my scheduled rounds did. Like a smoke alarm that beeps before you smell anything.
- analyze_policy_run.py — coverage and intervention counts broken down minute-by-minute, plus an automatic verdict against the three review criteria.
- eval_model_baseline.py — measures the model’s coverage in its native environment (5 episodes), giving us the reference bar.
- obs_parity_check.py — numerically compares the live observations against the training-environment raycast, so we know the model is seeing what it was trained to see.
5. Sources
- The three-criteria review framework (Aleks + independent reviewer): the coverage / escapes / guard-time acceptance bars.
- Active-mapping agent research that informed the world-state-based stuck detector (finding 5 of the night).
- The ArduPilot 4.8-dev parameter-rename discovery — surfaced by auditing the live autopilot tree against the config file (the day’s 🔴 finding).
- A parallel rl-lab parity track: a bridge map protocol document and an occupancy-map builder that matched 5/5 rl-lab fixtures to tight tolerance on the first run.
6. What’s next
- ActiveMapping v1.5: waiting on the next model from rl-lab, then wiring up its ROS2 harness.
- Distance-sensor → flight-controller path: a live smoke test plus a flight-capture video on the next stack restart. Groundwork already landed: a distance-sensor bridge node and the matching config for the wall-facing and downward range sensors, handed off to the firmware agent.
- A long run of configuration H to the full step horizon, to get clean statistics on how close the policy gets to the coverage cap without timeout noise muddying the picture.
- The sensor-monitor “fail loud instead of caching” change has already landed; the next step is verifying it on hardware.