claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

33 — Crash Recovery, exit(2), and the Relaunch-Climb Blocker

How a tumbling drone in SITL led us through crash-recovery escalation, an exit(2) safety net, and a relaunch-climb investigation.

stablesimulationupdated 2026-06-16T00:00:00.000ZClaudeDroneSimulationDevLog

What this is: A long debugging day in our simulator, told as a story. A training drone kept flipping over, the recovery machinery kept failing in a surprising way, and the real cause turned out to be nothing like what we suspected.

Why it’s here: Because the most useful lessons here are not the code patches — they are the detective work. We chased a “yaw bug” through five different fixes, all of which were validated and all of which missed the real culprit. The story shows how the simplest tool (watching a video) beat hours of log analysis.

Date: 2026-06-09 (a long afternoon-into-evening session) Ticket: SITL reinforcement-learning fine-tune sprint (D2 machine, several reboots)


Glossary

Before we dive in, a few terms in plain English.

  • SITL — “Software In The Loop.” Instead of flying a real drone, we run the exact flight-control software on a computer and feed it a simulated world. The drone “thinks” it is flying; it just lives inside a computer. This lets us crash a thousand times for free.
  • Gazebo — the 3D physics world the simulated drone flies in. It computes gravity, collisions with walls, and how the drone’s body tumbles. Think of it as the video-game engine behind the flight.
  • ArduPilot — the autopilot brain. It decides motor speeds, holds altitude, and keeps the drone level. The same software runs on real ClaudeDrone hardware.
  • mavros — the translator that lets our ROS2 code talk to ArduPilot. ROS2 speaks one language, the autopilot speaks another; mavros sits in the middle.
  • ROS2 (Jazzy) — the robotics middleware that ties all the software pieces together: sensors in, motor commands out. “Jazzy” is the specific version we use.
  • Crash recovery — what the system does when the drone falls over during training. Ideally it stands the drone back up and keeps going so a multi-hour training run doesn’t stop on the first stumble.
  • Soft reset vs. full relaunch — two recovery strengths. A soft reset nudges the drone back to a fresh state while the simulator keeps running. A full relaunch tears everything down — kills the simulator, the autopilot, the bridges — and boots the whole stack from scratch. Slow, but thorough.
  • Exit code — a single number a program returns when it stops, telling the outside world why it stopped. By convention 0 means “all good.” We deliberately picked exit code 2 to mean “something unrecoverable happened — a human needs to step in.”
  • Climb timeout — the drone is told “take off and reach 1.8 m.” If it hasn’t climbed by a deadline, that’s a climb timeout: it’s stuck on the floor pretending to fly.
  • Tumble — the drone loses control of its orientation and flips. Not a gentle lean — a genuine somersault.
  • Tilt angle — how far the drone is leaning from level, in degrees. 0° is perfectly flat; 90° is on its side.
  • Dataflash log — the autopilot’s onboard black-box recorder. After a crash we open it to replay exactly what the drone’s brain thought was happening, frame by frame.

1. The setup: a recovery system that escalates

We were deep into a fine-tuning run that needed the drone to survive thousands of training episodes back-to-back. A single crash shouldn’t end the run, so the recovery system was built to escalate: try the cheap fix first, and only reach for the expensive one if the cheap one fails.

We landed a chain of recovery fixes that day:

  1. A stuck climb is now treated as fatal. Previously, if the drone failed to take off, the code shrugged and “continued in hover” — which really meant grinding through worthless episodes with the drone sitting on the floor. The contrast was stark: the drone was logging dozens of episodes stuck near the ground for every handful that actually got airborne. Now a failed climb raises a real error and triggers recovery instead of being quietly ignored.

  2. A duration gate on the crash detector. A sharp acceleration can briefly tip the drone past 45° for a fraction of a second without it being a real crash. We added a rule: only latch a crash if the heavy tilt is held for at least a third of a second. This filters out the transient “false flip.”

  3. Crash → full relaunch. Once soft resets are exhausted, the system now does a complete teardown-and-reboot of the whole stack.

  4. Port cleanup on launch. Kill leftover processes so a relaunch doesn’t collide with a port that’s still held by a zombie from the last run.

2. The exit(2) safety net

There was a known bug in the relaunch path that we were explicitly told not to fix on the fly — it needed a proper investigation, not a quick patch under pressure.

So we built a guardrail instead. After the system exhausts its retry budget (three full takeoff attempts), rather than throwing a confusing error deep in the code, it now logs a clean message and exits with code 2, meaning “this needs a manual reboot.” The reasoning: with the new crash filter in place, genuine crashes during a long run should be rare. If one does slip through and recovery can’t handle it, that’s an exceptional event — better to stop cleanly and have a human reboot the machine than to thrash.

On the training side, the consuming code watches for that exit code, writes an “unrecoverable crash” marker, and shuts down tidily. No mysterious stack traces, no half-dead processes.

3. Smoke tests: the recovery machinery works, but…

We ran three quick “smoke” tests to confirm the new behavior. Here’s how they came out:

Smoke test Result
Crash escalation (after reboot) FAIL — all four fixes escalate correctly, but every full relaunch refused to climb (stuck near the floor)
exit(2) behavior PASS — episodes run, retries exhaust, the program exits cleanly with code 2, no stack trace
Angle limit PASS — the lean-angle limit loaded correctly, the drone climbed fine, and a burst of forward moves never falsely latched a crash (max lean ~1.2°)

That first FAIL is the important one. The escalation logic was perfect — but it kept escalating into a wall. Every time the system did a full relaunch, the drone simply would not take off again.

We also confirmed on the live autopilot that our lean-angle limit was genuinely being enforced (the autopilot reported the exact value we set, so the saved settings weren’t silently overriding it).

A small but real process lesson from this day: always launch the stack with logging enabled. One of our launches that day was started without it, so the simulator’s log froze hours earlier and we lost a chunk of evidence right when we needed it most.

4. Two real production crashes — both dying the same way

We kicked off two real training runs. Both died on the third episode, both for the same reason.

  • Run #1: The drone actually flew for a bit on the warm-started policy, then suffered a real flip — tilt shooting to 60° while it was up at 1.5 m. The next episode tried a full relaunch, hit the climb timeout three times in a row, and exited with code 2.

  • Run #2: (after a reboot) The first episode climbed fine, the second flipped, and the third tried three full relaunches — each one connected, reported the right flight mode, but sat stuck near the floor and never climbed. Three out of three identical failures, then exit(2).

The black-box replay of the first crash was revealing. The tilt came first — this was a genuine loss of control, not a sensor glitch. The drone’s lean climbed from 13° to 55° while altitude held steady, and only then did altitude collapse. That ordering matters: the fall was a consequence of the flip, not the cause. The roll and pitch were oscillating wildly — swinging between roughly −72° and +81°. That’s a tumble, not a steady lean. At its worst the drone reached 81° — nearly on its side.

We had bumped the lean-angle limit up that day hoping to give the policy more room. But in run #2 the lean still reached 57°, well past the limit. That confirmed something important: an angle limit governs commanded lean — when the autopilot deliberately tips the drone. It does nothing against an instability tumble, where the drone is losing control and flipping on its own. Different problem, wrong tool.

5. The open blocker: relaunch-climb

At this point we had a clean, reproducible blocker — serious enough to stop the whole long-run effort.

The symptom: A fresh boot climbs perfectly every time (the manual launch climbed, the smoke tests climbed, the first episode of both real runs climbed). But any full relaunch — kill the stack, boot it again — refuses to climb. The drone connects, accepts the takeoff command, and just… doesn’t lift. It reproduces after a single kill-and-relaunch cycle. It isn’t a graphics issue (the boot is fresh and the first episode climbed on the same GPU). It isn’t slow accumulation over time. Something in the relaunch path itself breaks the climb.

We listed candidate root causes but were forbidden from patching them blind — they needed a proper investigation first:

  1. The physics simulation doesn’t re-sync after a fast kill-and-respawn — the autopilot reconnects, but the physics never steps for this drone.
  2. The position estimator doesn’t converge in the short window before takeoff, so the altitude controller accepts the command but has no valid estimate to act on.
  3. A subtle difference between the manual launch (which climbed) and the recovery launch.
  4. The pause between kill and relaunch is too short, so the simulator’s network ports haven’t been released yet.

The strategic conclusion was uncomfortable: “exit and reboot on every crash” simply doesn’t scale to a long run. Real flips were happening too often. The real fix had to be the relaunch-climb path itself.

6. First resolution: the tumble’s root cause found

Then a key insight on the tumble (a separate problem from the relaunch blocker). Two independent investigations converged on the same culprit: an autopilot setting that makes it automatically steer its heading toward the next waypoint was fighting with our own explicit heading control.

We had two solid sources of proof — the black-box log and a live query of the autopilot both showed the setting at its default “auto-steer” value.

The mechanism: Our forward-motion command tells the drone to move via position targets. With auto-steer on, the autopilot tries to rotate the drone’s heading toward the target at the same time our code is commanding a fixed heading. Two heading controllers fighting each other → the orientation oscillates → loss of control → tumble. That perfectly explained the wild heading numbers we’d seen in the crash log.

The fix: turn auto-steer off entirely. Our code already steers the heading itself, in a closed loop — this setting should have been off from the start.

Validation looked clean:

Episode Actions Max tilt Max heading drift Tumble?
ep0 (full takeoff) forward × 8 1.1° 0.05° no
ep1 (soft reset) forward × 8 1.0° 0.04° no

Heading drift was essentially zero and tilt was around 1° instead of 57–81°. The tumble looked closed.

7. A second tumble mode appears

Except it wasn’t. A fresh run died on the very first episode — the heading fix held for the first move, but the second move (right after a couple of rotations) tumbled to 52°. The clue: the crash happened immediately after rotating.

We made the settle step attitude-aware. Previously, between actions the code just paused for a tenth of a second. But a rotation induces a roll/pitch wobble, and a tenth of a second doesn’t damp it out — so the next forward move stacked a lean on top of an already-wobbling drone. The new settle waits until the drone is genuinely level (under 5° tilt) before allowing the next move.

The smoke test was informative even though it failed:

Pair Pre-move tilt Result
0 0.15° OK
1 0.67° OK
2 0.72° OK
3 0.55° TUMBLE (ended at 63.5°)

So the rotation-transition mode was closed — the settle did its job, every move started from a level drone. But pair 3 tumbled anyway, starting from a perfectly stable 0.55°. The black-box showed the heading running away mid-move (from 55° to 124°), which then coupled into roll and pitch and flipped the drone. We checked our own code carefully and confirmed we were holding the commanded heading constant — so this runaway was happening on the autopilot side, not in our commands.

8. The heading-setpoint fix — closer, but still not it

Next diagnosis: the way we were sending position targets didn’t actually carry our intended heading through to the autopilot, so the autopilot was free-running the heading along the direction of travel.

We switched to a message type that carries an explicit heading, and we masked off everything except position and heading. Now the commanded heading did reach the autopilot and stayed constant during the move — a real, verifiable improvement.

And pair 3 still tumbled. The black-box now showed the commanded heading holding steady, but the drone’s actual heading was about 60° away from it. We command a heading, the autopilot sees the drone is 60° off, and it slews hard to correct mid-flight — which couples into roll and flips it.

This raised a worrying flag: the heading the policy sees during training comes from the same source that was 60° wrong. A 60° error there means a distorted view of the world during training, not just a tumble. That moved to the top of the priority list.

9. The pivot: it’s not a heading command at all

A fourth heading-related fix — fully separate the moving and the steering, so the autopilot never corrects heading while translating. Implemented and tested. Pair 3 tumbled again, this time to 65°.

This black-box reading was the turning point:

idx | DesRoll  Roll  | DesHeading  Heading
390 |  -2.2    -6.5  |   +52        +80      heading held constant (fix working)
391 |  -2.9   -27.8  |   +52        +98      Roll -28° when only -2.9° was commanded
392 |  +1.0   -70.5  |   +79       +140      Roll -70° physically -> tumble

The drone rolled to −70° when the autopilot had commanded only −2.9°. The autopilot never commanded that roll. So the tumble was a physical roll, not a heading command. All five heading-focused fixes had been refuted — every one of them failed on pair 3 in exactly the same way. The heading drift wasn’t the cause; it was a symptom of the flip.

We refined the theory: the drone’s estimated heading was drifting 30–40° over four rotations, which corrupted the autopilot’s internal leveling loop — it was applying roll and pitch corrections in the wrong frame of reference, causing cross-coupling and a physical roll-over. That elegantly explained why no heading command fix could help: the problem was in the estimate, not the command.

It was a good theory. It was also wrong.

10. The real resolution: watch the video

After five heading configurations, the decisive move was almost embarrassingly simple: run it with the graphical view on and record a video, instead of headless.

The video showed in seconds what five rounds of log analysis had missed. On pair 3, the drone crabs sideways toward a wall and slams into a corner — and that is the tumble. Its final position lined up exactly with a corner of the room. The root cause was a wall collision. Not a heading command, not an estimate drift. A wall.

A candid note on our own mistake: the “heading-estimate drift” theory was an over-confident reading of the data. The gap between commanded and actual heading in the black-box was an internal control error — not proof that the estimate disagreed with reality. We never actually compared the estimate against the simulator’s ground truth. And the log channel that should have shown ground truth was producing constant nonsense values in this configuration — itself a logging artifact. The real ground truth was the picture: the GIF.

The fix (test-only): make the forward move stop further from walls so it can’t push the drone into one. Crucially, this was a smoke-test artifact — the real policy’s sensor masking never drives those long moves into walls, so this only affected the test rig, not the shipped behavior. With that change, the full six-pair smoke test passed: max tilt 1.4°, zero tumbles.

The flight-path trace confirmed it: the drone does drift toward the corner over the run, but with the wall-stop in place the final move travels zero distance — it refuses to push into the wall — so there’s no collision, and tilt stays under 1.4° the whole flight.

Artifacts captured (a reminder that flight visuals belong in the dev-log):

  • Crash GIF: the drone colliding into the corner — the actual root cause.
  • Pass GIF: a clean flight after the wall-stop fix.
  • Two recorded videos: one of the crash, one of the clean pass.
  • The flight-path trace as JSON, plus the dataflash logs.

11. Where things stand

  • The tumble blocker is resolved — it was a smoke-test artifact (a too-aggressive test move running the drone into a wall). The real policy doesn’t fly into walls, so the long training run can proceed.
  • The relaunch-climb issue from section 5 remains open, but is now far less urgent: with the tumble gone, real crashes during training are rare, so recovery rarely triggers. If a long run ever does hit it, the section-5 hypotheses are waiting.
  • The heading fixes from sections 6–9 (turning off auto-steer, the attitude-aware settle, the explicit-heading message) all stand — each was validated against its own criterion. They were necessary improvements; they just weren’t the root cause of this tumble.

Process lessons

  • Keep the dev-log going always — don’t defer it to “after the sprint.” On a fast debugging day the journal slipped for hours and had to be reconstructed from logs and commits after the fact.
  • Always launch with logging enabled — otherwise the stack logs go missing exactly when you need them.
  • Write up an open problem as a standing document immediately (like section 5 above), not just as a note in someone’s head.
  • And when logs lie, look with your eyes. A few seconds of video ended a five-fix wild-goose chase.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR