What this is: A diagnosis of why our training drone keeps tipping over and crashing — written after pulling raw flight telemetry off the simulated aircraft many times per second and watching exactly what breaks first.
Why it’s here: For a long time we assumed the RL “brain” was making bad choices. It turns out the real culprit sits one layer below the brain, in the flight controller. This entry separates the two so future work targets the right thing.
Date: 2026-06-10 Ticket: rl-lab dev-log 31
Glossary
- Controller instability — when the low-level autopilot (the part that actually spins the motors to keep the drone level) starts overcorrecting and rocking back and forth instead of settling. Picture balancing a broomstick on your palm: small wobbles are fine, but if each correction is bigger than the last, the wobble grows until the broom falls. That growing wobble is instability.
- Reposition — the short hop the drone makes right after takeoff to get to its agreed starting point before the real run begins. Think of an actor walking to their mark before the scene starts.
- “Lunge” — the sudden, aggressive surge the drone makes when it’s handed a far-away target all at once. Instead of easing into the move, it leans hard and shoves itself forward, like a sprinter exploding off the blocks. Inside a small room, that lunge is exactly what gets you into a wall.
- Crash profile — the characteristic shape of how a fall actually unfolds over time: which thing breaks first, how fast, and what follows. Two crashes that look identical to the naked eye can have very different profiles when you read the numbers.
- PPO — the reinforcement-learning method we use to train the drone’s high-level decision-making (the “brain”). Important: PPO chooses where to go; it does not directly drive the motors. That’s the autopilot’s job.
- tilt — how far the drone is leaning from level. A few degrees is normal cruising. Past roughly 50° there’s no recovery — that’s a flip and a fall.
- carrot — a “carrot on a stick” target. Instead of telling the drone “go to that far point now,” we dangle the goal a short distance ahead and keep moving it forward, so the drone follows smoothly.
- autopilot (ArduPilot) — the low-level program that holds the drone steady at a point. It’s the reflex layer, separate from the RL brain.
- policy (the RL “brain”) — the trained model that picks an action (which way to fly) but never touches the motors directly.
1. What I wanted
I wanted a clear, evidence-based answer to a question that had been nagging at us: when the drone tips over and falls, is the brain making a bad decision, or is something underneath it failing?
This matters enormously for where we spend effort. If the brain is choosing badly, the answer is “train more / train better.” If the layer beneath the brain is unstable, no amount of training will help — we’d be teaching a smart pilot to fly a plane whose controls fight back. I wanted to settle that with numbers, not hunches.
2. What I tried
I instrumented the simulated drone and recorded its real flight telemetry — tilt and altitude — sampled many times per second (roughly 10–50 readings every second). That sampling rate is the whole point: a crash that takes under a second to unfold is invisible if you only glance at the drone, but at this resolution you can watch the failure frame by frame, like slowing a video down until you can see the exact instant something snaps.
I then compared three situations side by side: a normal smooth move, the suspicious post-takeoff reposition, and a full crash.
3. What happened
The data was clearer than I expected. Three findings, in order.
Finding 1 — Normal movement is flawless. When the target is dangled smoothly in front of the drone (the “carrot” approach), the drone holds its altitude rock-steady and never leans more than about 2°. There is simply no problem here. This is the baseline that proves the drone can fly clean.
Finding 2 — The “weird start” is real. During reposition, the drone is handed a target on the far side of the room all at once, with no carrot to ease it in. The autopilot responds with a big lean immediately, and that lean overshoots into a rock — up to 16°, two full swings, before it catches itself. Out in open space it survives this. Next to a wall, those same two swings would put it into the wall.
Finding 3 — The crash profile: rock first, then a sudden break. In a failed move, the failure has two distinct phases:
| Phase | What happens | Duration |
|---|---|---|
| Build-up | Tilt rocks with growing amplitude: 2° → 6° → 11° → 14°. Altitude holds steady the whole time. | ~7 seconds |
| Break | Tilt jumps 14° → 64°, the drone flips, and only then does altitude collapse: 1.85 m → 0.23 m in about a second — nearly free-fall. | ~0.3 seconds |
The most useful takeaway answers a question we’d been asking directly: how does the altitude drop — does it sag gradually? No. Altitude stays level and then drops like a stone, and it does so after the flip, not before. The tilt is what fails first (the rock builds, then snaps); the altitude loss is just the consequence — a drone that’s upside down can’t make lift.
So the chain of causation is unambiguous: growing tilt oscillation → flip → altitude collapse. The brain isn’t in this chain at all.
4. The conclusion
The root cause is the autopilot oscillating when it’s handed sharp or diagonal targets in one shot. This is not something training can fix — it lives at the flight-control layer, below the brain.
The right fix is to change how movement commands are issued: instead of “fly to this point,” the sim should drive the drone with smooth velocity commands expressed in the drone’s own frame of reference. That keeps motion gentle and stops the oscillation from ever starting. Reposition needs the same smoothing — right now it slams the target down all at once, which is exactly the source of that post-takeoff lunge.
Once that’s in place, the RL brain inherits a stable drone. Then PPO can spend its capacity learning the high-level job — where to look, where to go — instead of wasting it trying to learn “how not to somersault.”
5. Sources
- Raw flight telemetry (tilt + altitude) recorded directly off the simulated drone at ~10–50 Hz, across normal moves, reposition, and full crashes.
- Direct comparison of the three flight situations described above.
6. What’s next
- Move command issuance to smooth velocity commands in the drone’s own frame, so the autopilot is never asked to make a sudden corner.
- Apply the same smoothing to the reposition hop to kill the post-takeoff lunge.
- Re-run with the stabilized controller and confirm the tilt-oscillation signature is gone from the telemetry before asking the brain to learn anything new.