claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

34 — RCA: The Maintenance Stream That Hijacked Re-Takeoff

How a stale 10 Hz setpoint stream silently overrode NAV_TAKEOFF after relaunch, and the five-line fix that made crash recovery robust.

stablesimulationupdated 2026-06-16T00:00:00.000ZClaudeDroneSimulationDevLog

What this is: A root-cause investigation into why our simulated drone refused to climb when it tried to take off a second time after a crash — and the small fix that finally made recovery reliable. Why it’s here: It is a clean example of how a “red herring” can mislead an entire debugging session, and how a disciplined, evidence-first method (measure two things at once, then compare) cut through it.

Date: 2026-06-10 (late night, ~02:00–02:40) Ticket: Problem B — re-takeoff after relaunch does not climb


Glossary

  • RCA (root-cause analysis): Instead of patching the symptom, you keep asking “why?” until you reach the one thing that, if changed, makes the whole problem disappear. Like a doctor finding the infection rather than just lowering the fever.
  • Maintenance stream: A steady background heartbeat. Our flight controller likes to be told “go to this position” many times per second. The maintenance stream is the component that repeats that command at 10 Hz (ten times every second) so the drone stays put where we asked.
  • Re-takeoff: Taking off a second time, on a stack that has already flown once and then been reset after a crash. As we will see, “the first takeoff” and “the second takeoff” turned out to be two very different stories.
  • Setpoint: A target — “be at position X.” If you keep shouting an old target at the drone, it will obediently try to stay at the old spot even when you want it to do something new.
  • NAV_TAKEOFF: The dedicated “lift off and climb to this altitude” command in the flight controller. It is a higher-level intent than a raw position setpoint.
  • SITL (Software In The Loop): Running the real flight-controller firmware as a program on a PC, with no physical drone. It thinks it is flying; we feed it simulated sensors.
  • Gazebo: The physics simulator that provides the “real” world — gravity, motors, lift. It is the ground truth of where the drone actually is.
  • mavros: The bridge that lets ROS2 talk to the flight controller. Commands and telemetry pass through it.
  • EKF (the estimator): The flight controller’s internal best-guess of where it is, fused from its sensors. It can disagree with the true physics if something desyncs.

1. The starting hypothesis (which was wrong)

This investigation picks up from dev-log 33, where a crash-recovery run kept failing: after a wall crash, the agent would try to relaunch the simulator three times in a row, the drone would sit stuck at about 0.21 m off the ground, and the process would give up with exit code 2.

In dev-log 33 the failure was pinned on the simulation infrastructure — guesses included the physics link not re-syncing after relaunch, the estimator having no position fix, instance mix-ups, and timing/port issues. My own first read of the logs blamed a physics-link desync: the idea that the flight controller’s estimate was “flying ahead” of the true Gazebo physics.

The honest answer at the time was: reboot the machine and it works again. That “fix” is exactly the kind of thing that should make you suspicious.

2. A discriminating test: measure both altitudes at once

The method here is simple and worth stealing for your own debugging: don’t measure one number, measure two, and watch whether they agree.

We built a small smoke test that logs two height trajectories simultaneously:

  • SITL-z — the altitude the flight controller believes it is at (its estimator output, read via mavros).
  • gz-z — the altitude the drone actually is at, read straight from Gazebo physics (ground truth).

If those two numbers diverge, the problem is a desync between belief and reality. If they move together, desync is ruled out.

The test ran in two phases: Phase A was a clean, fresh takeoff (the baseline). Phase B restarted the flight controller and attempted a re-takeoff, escalating internally up to three full relaunches if it kept failing.

Result — the failure reproduced, deterministically

Phase SITL-z max gz-z max Reading
A — baseline (fresh) 2.0 / 1.83 1.99 / 1.83 Both climb; the ground-truth probe is valid (gz == sitl)
B — restart + 3× full relaunch 0.21 0.22 Both stuck at spawn height; armed, NAV_TAKEOFF accepted, no climb → exit(2)

This is the key insight. In the failing case, both altitudes were stuck together — the believed height and the true height both sat at ground level. That immediately rules out a desync: the flight controller was not “flying in its head” while physics lagged. Its estimate was honestly stuck too.

So two ideas died here at once: my desync theory, and the dev-log 33 hope that “a full relaunch with a fresh simulator fixes it.” A full relaunch did not fix it either.

3. The decisive probe: command versus effect

After the failing run gave up, the whole software stack was still alive. So instead of throwing it away, we poked it.

We sent a manual takeoff command straight through mavros to that same, supposedly broken stack: switch to guided mode, arm, take off.

motor PWM 1500–1640 (real thrust!) · gz_z 0.22 → 1.98 m · rotors spinning OK

The drone climbed to 1.98 m. On the exact stack that had just “failed.”

That is conclusive. The simulator, the physics link, the motors, the recovery path — every piece of infrastructure was healthy. Problem B was not an infrastructure bug at all. The bug had to live in the takeoff code path itself, specifically after a relaunch.

4. The root cause (airtight)

Here is what was actually happening.

The action executor runs that constant 10 Hz maintenance timer described in the glossary. Ten times a second it publishes a position setpoint to the flight controller — “stay here.” Its guard logic only stayed silent in a few specific situations. Crucially, there was no guard for the takeoff window.

The code even had a comment admitting the danger: take off with NAV_TAKEOFF and climb without the setpoint stream running, because the stream overrides NAV_TAKEOFF. The author knew the two commands fought each other.

Now compare the two takeoffs:

  • The first takeoff (baseline): The drone has no target yet, so the maintenance stream has nothing to send and stays quiet. NAV_TAKEOFF runs unopposed, the drone climbs, and only afterward does the code set a target to hold.
  • The re-takeoff (after relaunch): The old target from the previous episode is still alive — the reset and relaunch never cleared it. So the moment the new takeoff begins, the maintenance stream is already shouting the stale old setpoint at 10 Hz. That stale command overrides the fresh NAV_TAKEOFF. The drone dutifully holds its old position (~spawn height, 0.2 m), the climb times out, recovery kicks in, and the process exits.

And the reboot “fix”? A pure red herring — the very one that misled dev-log 33. Rebooting restarted the Python training process, which created a fresh communication object with no leftover target. So the first takeoff worked again. The reboot wasn’t healing the stack; it was just wiping the stale target. That accidental cleanup produced the false conclusion that this was a relaunch/physics bug.

5. The fix — applied and validated

The fix is almost embarrassingly small once the root cause is clear: before the climb, clear the stale target so the maintenance stream has nothing to shout.

We added a clear_target() method on the action executor (it resets the held target and its path segment) and called it inside the full-takeoff routine before the climb begins. With the target cleared, the maintenance stream goes quiet during the takeoff window — exactly recreating the baseline condition that always worked. After the climb finishes, the existing code re-establishes the target as normal. Five lines, and it does not change the in-flight training behavior at all — only the takeoff window is touched.

Validation — re-running the exact scenario that failed

Phase SITL-z max gz-z max Before fix After fix
B — restart + re-takeoff 1.92 2.02 z=0.21 → 3× full relaunch → exit(2) climb to 1.86 → stable hover, no escalation

The communication log now reads cleanly: “NAV_TAKEOFF 2.0 m, climbing… climb done z=1.86 m → hover stable ≥5 s.” Recovery after a flight-controller restart is now robust — meaning a long training run can survive crashes without a reboot every time, which was the whole point.

One honest caveat for whoever comes next: the ground-truth altitude probe occasionally returns NaN under heavy load (in Phase A it once falsely reported no value). That probe deserves a retry. It does not affect the conclusion here — Phase B is the decisive phase, and there both altitudes read ~2.0 m.

6. Lessons worth keeping

  1. When a reboot “fixes” something, distrust it. A reboot resets a hundred things at once. It is the opposite of a controlled experiment, and it will happily hide the real cause.
  2. Measure two signals and compare them. The single move that broke this case open was logging believed-altitude and true-altitude together. Their agreement instantly killed the most popular theory.
  3. Don’t throw away a broken stack — interrogate it. The manual-takeoff probe on the dead stack proved the infrastructure was fine and redirected the entire search.
  4. State that outlives an episode is a classic trap. The bug was a leftover target surviving a reset. Whenever you reset, ask what didn’t get cleared.

Artifacts

  • Discriminating smoke test (dual-altitude: ground truth vs. flight-controller estimate).
  • Recorded altitude tracks and runtime logs for both the reproduction and the passing run after the fix.
  • Key code: the maintenance-stream timer in the action executor, and the full-takeoff / climb routines in the communication layer.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR