claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

29 — Why SITL Fine-Tuning Kept Crashing the Drone

Three SITL fine-tuning runs all tumbled the drone within ~20 steps. The real culprit wasn't sensors or restarts — it was the policy itself.

stablerl-labupdated 2026-06-16T00:00:00.000ZClaudeDroneRLDevLog

What this is: A field note from trying to take our trained 2D drone “brain” and fine-tune it inside a realistic flight simulator. Three attempts, three crashes — and the story of how we traced the failure to its real source.

Why it’s here: Because the lesson is counter-intuitive. We were sure the simulator was at fault. It turned out the drone’s own policy was throwing it to the ground. That’s a finding worth writing down.

Date: 2026-06-09 · Machine: D2 (GPU) Ticket: TASK-RL-SITL-FT-1


Glossary

If you’re new to the vocabulary, read this first. Everything below leans on these terms.

  • Policy (the “brain”) — the trained model that decides, moment by moment, what the drone should do: where to fly, where to look. We already have a good one, trained in a flat 2D world.
  • SITL (software-in-the-loop) — a realistic simulator that runs the actual drone flight controller software (ArduPilot) inside a 3D physics engine (Gazebo). Think of it as a flight rehearsal where the real autopilot brain is present, just the body and the world are virtual. It’s far closer to reality than a 2D toy world.
  • Fine-tune — take a model that already works and nudge it to fit new conditions, instead of training from scratch. Like a pianist who can already play, sitting down at an unfamiliar piano and adjusting their touch — not relearning music.
  • Warm-start — starting the fine-tune from the existing model’s weights rather than from a blank slate.
  • Tilt — how far the drone leans. Lean too far and it flips over and falls.
  • Tumble — the drone has lost control and is spinning chaotically. In plain terms: it’s crashing.
  • Blocker — something that stops progress cold. Here, the thing blocking fine-tuning was the tumble itself.
  • PPO — the reinforcement-learning method we use to train and fine-tune the policy. It’s the algorithm doing the “nudging” during fine-tune.
  • Stack relaunch — restarting the whole simulator from scratch (kill everything, bring it back up).
  • exit(2) — our deliberate emergency stop. Instead of hanging forever, the program says “that’s it, this machine needs a restart,” and exits cleanly.

1. What I wanted

Take our working 2D drone policy and, over 50,000 training steps inside the realistic SITL simulator, teach it to fly well there. This is the bridge toward a real, physical drone — the 2D world is a sketch, SITL is the dress rehearsal, hardware is opening night.

There was a hidden assumption baked into the plan: “the model already knows how to fly, we just need to polish it for the new stage.” Hold onto that sentence. It turned out to be false, and that falseness is the whole story.

2. What I tried

I ran the fine-tune in three separate attempts. Between each one, the simulation agent fixed different bugs in the stack. Each attempt failed in its own way — and the pattern in how they failed is the interesting part.

  1. Run 1 — the drone barely got off the ground. After landing it couldn’t take off again — it just hovered stuck at about 21 cm above the floor, and did that 120 times in a row. Every “episode” was useless: imagine a learner driver whose car won’t pull away from the curb, so there’s nothing to learn from.
  2. Run 2 — it actually took off, flew for roughly 19 steps, then tilted to 60° and fell. The simulator restarted itself after the crash, but again couldn’t get airborne → emergency stop.
  3. Run 3 (now with two fresh anti-fall fixes in place) — took off, flew only 7 steps, tilted to 57° and fell. Restart → still won’t take off → emergency stop.

3. What happened (the real finding)

The problem turned out to have two layers, stacked on top of each other.

Layer 1 (simulator side — not my zone). After any simulator restart, the drone won’t take off — it gets stuck at that same 21 cm above the floor. That means auto-recovery after a crash simply doesn’t work: any fall becomes a dead end. You crash once, and the rehearsal is over.

Layer 2 (new, and the heart of it — my zone). The drone wasn’t falling because of a glitchy sensor. It was falling for real. Our 2D policy issues commands that make the drone in SITL lose control and tumble within 7–19 steps. And here’s the part that matters most: both fresh fixes failed to help.

  • We added a tilt limit (25°) — yet the drone still reached 57° and even 81°. Why? The limit holds the requested tilt. But once the drone loses control, physics flips it far past anything we asked for. It’s like capping how hard a driver may turn the wheel — useless once the car is already skidding.
  • The second fix simply detects the fall correctly. But the fall is genuine. We don’t need to detect it better — we need to prevent it from happening at all.

The takeaway: we used to blame sensors, parameters, or the restart logic. Now it’s clear the policy itself is the cause — its commands knock the drone down. So “reboot the machine and try again” is pointless: the drone will just fall again in about 10 steps. You can mop the floor all you like, but the tap is still running.

One genuinely good outcome: our emergency stop worked flawlessly. Every single time, the program shut down honestly and cleanly, breaking nothing.

Run Took off? Steps before fall Tilt reached Outcome
1 barely (stuck at ~21 cm) — — 120 useless episodes
2 yes ~19 60° crash → no re-takeoff → emergency stop
3 yes 7 57° (peak 81°) crash → no re-takeoff → emergency stop

4. Sources

  • Source dev-log: rl-lab/docs/dev-log/29-sitl-finetune-50k-tumble-blocker.md
  • Ticket: TASK-RL-SITL-FT-1
  • Stack: ArduPilot SITL + Gazebo, driven via mavros under ROS2 (Jazzy)

5. What’s next

I’m waiting on a decision from Aleks / the simulation agent. There are three paths on the table:

  • (A — my first recommendation) Open up the flight recording like a black box and find exactly which command throws the drone down. My strong guess: we need to limit how fast commands are allowed to change — so the drone can’t jerk violently — rather than capping the tilt angle. Smooth the hand on the controls, not just the maximum angle.
  • (B) The simulation agent fixes the restart so that recovery after a crash actually works again.
  • © On our side, make the policy “gentler” — soften the strength and abruptness of its commands.

The honest conclusion of this run: the bridge to hardware needs the policy itself to be trained for stability under real physics, not just patched at the edges.

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR