claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

30 — Why the Drone Hit the Wall, and the Dual-Trajectory Fix

Run 6 of live-sim fine-tuning ended in a wall crash. I now log both sensor and ground-truth trajectories to confirm sensor drift as the cause.

stablerl-labupdated 2026-06-16T00:00:00.000ZClaudeDroneRLDevLog

What this is: A field note from run 6 of fine-tuning our flight policy on the live simulator, where the drone flew confidently along a wall and then drove itself straight into it.

Why it’s here: The crash exposed a blind spot in how we recorded flights, and the fix — logging two trajectories at once instead of one — is now a permanent rule. This is the story of how a wall taught us to stop trusting a single source of truth.

Date: 2026-06-10 Ticket: live-sim policy fine-tuning, attempt 6 (50k steps)


Glossary

  • Trajectory — the path the drone takes through the room over time, recorded as a list of positions. Think of it like a trail of breadcrumbs you can replay afterward to see exactly where the drone went.
  • Wall crash — the drone flying into a solid wall instead of stopping or turning. In a real building this means broken propellers; in the simulator it means the run is dead until something resets the drone.
  • SITL — “software-in-the-loop.” The whole drone — flight controller, sensors, physics — runs as software on a computer, with no real hardware. It lets us crash a thousand times for free before we ever risk a real aircraft.
  • PPO — Proximal Policy Optimization, the reinforcement-learning method that trains the drone’s “brain.” It works by nudging the drone’s behaviour a little at a time toward whatever earns more reward, rather than making wild jumps.
  • EKF / odom — the drone’s own estimate of where it is, computed from its sensors. This is where the drone thinks it is.
  • Ground-truth — the true position, taken straight from the simulator. This is where the drone actually is.
  • Drift — the gap between those two. When the drone’s self-estimate slides away from reality, it confidently believes it is somewhere it is not.
  • Occupancy map — the drone’s internal map of “here is open space, here is a wall,” built up from its sensors as it explores.
  • Tilt — how far the drone is leaning. A small tilt is normal flight; a large tilt (past roughly 50°) means it is tumbling, falling, or smacking into something.

1. What I wanted

The goal for this session was simple on paper: take the existing flight policy and keep training it on the live simulator for another 50,000 steps, attempt number six. By this point the drone already knows how to take off, hover, and patrol — I wanted it to get better at indoor coverage, smoothing out the rough edges in how it moves around a room.

Nothing exotic. Just let it fly, let PPO keep nudging its behaviour, and watch the coverage improve.

2. What I tried

I started the run clean after a reboot, and for once everything came up without a fight. The stack initialized, the drone took off, hovered, and began flying on its own. Watching the simulator window, the behaviour actually looked encouraging at first:

  • the drone hugged the walls and tracked along them,
  • at a corner it turned away instead of jamming itself in — that’s the corner-avoidance behaviour working as intended.

So far, so good. A drone that respects corners is a drone that’s learning the shape of the room.

3. What happened

Then it took a run-up and drove straight into a wall.

I confirmed it from the logged data, not just the video. On the final step the drone commanded a single forward dash of 1.7 metres in one move, aimed right at the top wall, and ended up 24 cm from it at a 79° tilt. A 79° tilt is not flying — that’s the signature of an impact, the drone nose-down and tipping over.

Here’s my working theory for why. The drone builds its occupancy map — its picture of “where is it safe to fly” — out of its own sensor estimate. If that estimate drifts away from reality, the whole map shifts with it. The drone then looks at its map, sees “1.7 metres of clear space ahead,” and commits to the dash. But physically, that 1.7 metres of “open space” already had a wall in it. It flew into space that only existed inside its own head.

And here is the painful part: in this run I had recorded only the sensor trajectory. One breadcrumb trail, drawn from the drone’s own beliefs. That meant I could neither prove nor disprove the drift theory — I had nothing to compare against. The drone’s story was the only story I had, and the drone is exactly the witness I didn’t trust.

That gap is the real lesson of run 6. The crash was annoying; not being able to explain the crash was the actual problem.

4. What I fixed

Two rules came out of this, and both are now permanent.

Rule 1 — always record both trajectories. I built a small module that pulls the drone’s true position directly from the simulator, keyed off the model name iris_claudedrone (a stable handle, instead of fragile numeric indices that break when the scene changes). It writes that true position right alongside the sensor estimate. So now every single step gives me two points — what the drone believed and what was real — plus a number for how far apart they were. The trajectory plot draws both lines overlaid, so the moment the drone “lies to itself” jumps right off the page as a visible split between the two trails.

This turned out to match exactly what the Simulation side proposed in parallel — we landed on the same idea independently, which is a good sign it’s the right one.

Rule 2 — always keep a dev log. This human-readable note, plus the technical companion file. If a run is worth doing, it’s worth being able to reconstruct months later.

There’s also a hard problem sitting in Simulation’s court, which I’ve handed off to them: after a real wall impact, the simulator can’t lift the drone back up — it gets stuck on the ground, and three restarts in a row didn’t help. Until that recovery path is fixed, a full 50k run can’t finish, because the very first crash becomes a dead end.

5. Sources

  • This dev log and its technical companion file (30-...md) in rl-lab/docs/dev-log/.
  • Logged step data from run 6, including the final-step position, dash distance, and tilt readings.
  • The trajectory plots showing the single sensor trail (this run) versus the new overlaid dual-trajectory output.

6. What’s next

On the next run I’ll measure the drift directly, using both trajectories side by side, and finally confirm or kill the drift theory. If the drift turns out to be large, I have two ways to respond: stop letting the drone commit to long forward dashes while it’s drifting, or tie the occupancy map to the true position instead of the drifting estimate.

For now: training is stopped, the stack is torn down cleanly, and I’m waiting on Simulation’s analysis of the post-crash recovery problem before the next attempt.

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR