A reinforcement-learning drone flips over and falls out of the sky. The first instinct, every time, is to blame the brain: the learned policy must have made a stupid decision. Train it more. Train it better. Add more data.
That instinct is usually wrong. More often than not, the policy decided correctly and the layer underneath it — the flight controller — is what actually broke. Spending weeks retraining a policy to fix a controller problem is like teaching a driver to steer more carefully when the real issue is that the car’s steering rack fights back. No amount of driver skill fixes a car whose controls overcorrect.
This article is about how to tell the two apart, and why getting it right saves enormous amounts of wasted effort.
A two-layer stack
Every autonomous drone runs at least two distinct layers of control, and they have very different jobs.
The policy (in our stack, a PPO-trained navigation model) is the high-level decision-maker. It looks at the world and decides where to go: turn left, advance toward that gap, hold position. It never touches the motors. Think of it as a passenger in the back seat giving directions — “next right, then straight on.”
The autopilot (for us, ArduPilot) is the low-level reflex layer. It takes the policy’s “go there” and turns it into the rapid, continuous motor adjustments that keep the aircraft level and moving. It is the driver’s hands on the wheel and feet on the pedals, executing the directions moment to moment.
When a drone crashes, the failure happened in one of these layers. The whole diagnostic art is figuring out which — and the layers leave very different fingerprints.
Reading the fingerprints
The trap is that to the naked eye a crash is just a crash: the drone tips and falls. You cannot tell layers apart by watching. You have to read telemetry — tilt and altitude sampled many times per second (tens of readings per second), fast enough to slow the failure down frame by frame and watch which thing breaks first.
That ordering is everything. A bad decision shows up as a sensible-looking control loop driving the drone somewhere it shouldn’t be — smooth flight, wrong destination, then a collision. A control failure shows up as the autopilot itself becoming unstable: small corrections that grow into bigger ones instead of settling, like trying to balance a broomstick where each catch is more violent than the last, until it falls.
We caught exactly this signature in our own reposition-lunge crash profile. When the drone was eased toward a target gradually — a moving “carrot” goal dangled just ahead — it held altitude rock-steady and never leaned more than about 2°. Clean flight. The brain was fine. But when it was handed a far target all at once, with no easing, the autopilot answered with a hard lean that overshot into a rock: up to 16°, swinging back and forth before it caught itself.
The full crash made the causation unmistakable. The failure had two phases. First a long build-up — roughly seven seconds of tilt oscillation growing 2° → 6° → 11° → 14°, with altitude holding perfectly level the entire time. Then the break: tilt jumped to 64° in a fraction of a second, the drone flipped, and only then did altitude collapse, falling from 1.85 m to 0.23 m in about a second — near free-fall.
Notice the order. Altitude did not sag gradually as a bad-decision story would predict. It stayed flat and then dropped like a stone, after the flip. The tilt failed first; the altitude loss was just the consequence, because an upside-down drone makes no lift. The chain reads: growing tilt oscillation → flip → altitude collapse. The policy never appears in that chain at all.
The root cause was the autopilot oscillating whenever it was handed a sharp, one-shot target. No amount of policy training touches that. The fix lived entirely at the control layer: issue smooth velocity commands in the drone’s own frame instead of slamming a distant waypoint down, and apply the same smoothing to the post-takeoff reposition hop so it never lunges in the first place.
Takeoff: same lesson, different bug
Crashes that happen before the policy is even in charge are the clearest proof of all. After a machine reboot, our drone began crashing on every single takeoff with an attitude-error abort less than a meter off the ground — and the policy hadn’t issued a single navigation command yet. Documented in simulation dev-log 29, the causes were pure flight-control timing and configuration:
- The drone armed before its state estimator (the EKF) had settled after the slower post-reboot sensor startup, so it lifted off with inconsistent gyros and immediately tipped. The fix was a fixed settling pause before arming — give the math time to agree with itself.
- The takeoff hover target was pinned to the world origin instead of the drone’s actual liftoff spot, so it banked toward origin during climb, leaned to nearly 47°, and sank. The fix was to read the real takeoff position and hold that.
Both were validated across six clean takeoffs. Neither had anything to do with the learned policy. If you had tried to “train away” these crashes, you’d have trained forever.
A diagnosis checklist
When your RL drone goes down, work the layers before touching the policy:
- Get telemetry, fast. Tilt and altitude at tens of Hz. You cannot diagnose what you only saw at human frame rate.
- Find what breaks first. If tilt destabilizes before altitude moves, suspect the controller. If the drone flew smoothly into a wall or wrong spot, suspect the decision.
- Look for growing oscillation. Corrections that escalate instead of settling are a control-instability signature, not a planning one.
- Check altitude shape. Sudden post-flip free-fall = consequence of a flip (control). Gradual sag = something else.
- Test with smooth vs. sharp targets. Clean on a gentle “carrot,” unstable on a one-shot target? The controller can’t handle sharp commands — that’s not the brain.
- Rule out the pre-policy phase. Takeoff, arming, heading hold, reposition all run before navigation. A crash there is never the policy’s fault.
- Only after all that, question the decision. If flight was stable and the destination was wrong, now it’s the policy’s turn.
Why it matters
Getting this right is not academic. Every hour spent retraining a policy to fix a control bug is an hour wasted, and worse, it can bury the real problem under noise. A policy can only ever be as good as the aircraft executing its commands. Stabilize the controller first, and the policy inherits a drone it can actually fly — free to spend its capacity on the real job of where to go, not on the impossible job of learning how not to somersault.
When the drone falls, read the telemetry before you blame the brain. The autopilot is the usual suspect.