claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

21a — The Model Beat the Classic, We Found Its Weak Spot, and Why We Froze Training

Our RL drone learned indoor coverage 5-7x faster than the 1997 classic, but spins in place on complex layouts. Why we paused to lock parity first.

stablerl-labupdated 2026-06-16T00:00:00.000ZClaudeDroneRLDevLog

What this is: A plain-language report on a training milestone — the model hit its target numbers, beat the classical baseline, and revealed exactly one weak spot. It also explains a deliberate decision to stop and not fix that weak spot yet.

Why it’s here: Sometimes the most important engineering decision is the one where you put your hands in your pockets and refuse to write code. This entry documents one of those moments, and the long history of pain that justifies it.

Date: 2026-06-07 Ticket: AM3 — train, oscillation, parity freeze


Glossary

A few terms appear throughout. If you only read one section, read this one.

  • Indoor coverage — the drone’s job in this task: fly through an unfamiliar set of rooms and “paint” as much of the floor map as possible, as fast as possible. Think of a robot vacuum that needs to see every corner, not clean it.
  • The classic (1997) — a textbook, hand-written exploration algorithm from the academic literature. It is reliable but plodding — the “diligent student” who always gets the right answer eventually. It is our baseline to beat.
  • PPO — the reinforcement-learning method we use to train the model. You don’t need the math: think of it as a way of letting the drone try millions of moves in a simulator and gradually keeping the habits that worked.
  • Training oscillation — when a trained model gets stuck doing the same back-and-forth move forever, like a person frozen between two doors, stepping left, then right, then left again, unable to commit. It’s not a crash; the model is “thinking” — it just never decides.
  • The bridge — code on the simulation side that translates the model’s chosen action (“turn left”, “go forward”) into real commands for the drone, and assembles the model’s view of the world in return. It is the interpreter standing between the brain and the body.
  • Parity (parity freeze) — the rule that what the model sees during training must match, exactly, what it sees in the real run — down to the last detail. A “parity freeze” means we stop training until both sides agree on that exact picture. This is the religion of this sprint.
  • Gazebo — the 3D physics simulator where the test rooms live. The drone flies in Gazebo before it ever flies for real.
  • SITL — software-in-the-loop: running the flight controller as software, so the whole stack can be tested without hardware.

1. The headline: the sprint’s target numbers are in the bag

The model trained quickly — about six minutes on an RTX 5070 GPU. (This particular task learns startlingly fast; a fresh task picking up the skill in minutes still feels like a small miracle.)

We tested it on 6×6 rooms — the exact same rooms that sit in our Gazebo world. Here is how it stacks up:

Who How much of the map Steps to reach 80%
Sprint target ≥ 80% ≤ 800
The classic (1997) 85% 206–325
Our model 89% 42

The model maps a room 5 to 7 times faster than the classic algorithm.

On an empty room it finishes in 19 steps: one sweeping look around from a good vantage point, a short flight to a new spot, one more look, done. That is exactly the behavior Aleks described in words long ago — “it stops, scans, flies over, scans again.”

2. The weak spot — found, reproduced, and fully understood

On hard “apartment” layouts (four to six rooms), the model gets stuck in 19 out of 50 episodes. That’s not a rounding error; it’s a real failure mode we have to respect.

I opened up one of those stuck episodes frame by frame. After step 73, the model spent 926 steps in a row spinning in place — left, right, left, right — and never escaped.

Here is why, in plain terms.

The hint the model uses to decide “where is the unexplored area?” is computed in a straight line — straight through walls. Picture a compass that points at buried treasure, but it points through the wall, not toward the door you’d actually use to reach it.

So the model turns to face the treasure. It can’t go that way — there’s a wall. From the new angle, the compass now points at a different patch of treasure. The model turns toward that one. New angle, new pointing, back to the first. And so on, forever. It twitches between two directions like a donkey starving between two equally tempting bales of hay.

This is a classic case of training oscillation: the model isn’t broken, it’s indecisive, because the information it’s given is subtly misleading.

The fix is conceptually simple: make the compass point not “as the crow flies” but “the way you’d actually walk” — the first step of a real route around the obstacle, the same way the classic algorithm navigates. One change, roughly an hour of work, then a retrain.

3. Why I did not run that fix — the stop call was right

Here’s the trap, and it’s a deep one.

That “compass” will not be computed by my training simulator when the drone flies for real. It will be computed by the bridge, on the simulation side, using their map and their algorithm.

If their way of computing it differs from mine even slightly — a different route around an obstacle, a different rounding rule, a different order of checking neighboring cells — then the model in Gazebo gets different compass readings than the ones it trained on. And it happens silently. No errors in the logs. The model just quietly gets dumber.

We have lived this exact nightmare before. A distance sensor (the TF-Luna) was once mounted rotated 90° from what the model assumed. Everything ran, nothing crashed, and we hunted that ghost for a month and a half.

So fixing the compass first and aligning it later would be building on sand. The right order is: agree on the exact picture, then train against it.

4. What I did instead of training

Rather than rush a fix that could become a month-long ghost hunt, I spent the time making sure both sides will produce a byte-for-byte identical world view.

  1. I read every relevant simulation document. The key finding: the bridge currently has no map-building code at all — it only tracks “where I have been” for the older model. That means there is no existing implementation to conflict with. We get to agree on the design before either side writes the code. This is the ideal moment — the cheapest possible time to lock parity.

  2. I sent simulation a structured questionnaire (logged in the team handoff at 11:30) covering six areas: how to build the map from sensor rays, what counts as a frontier between known and unknown space, exactly how to navigate around obstacles (down to the order in which neighboring cells are checked — that detail genuinely changes the route), how to normalize the numbers, and how to communicate “forbidden moves.” The format was deliberately collaborative: “here is my proposed canon — confirm it or push back,” so they aren’t designing from a blank page.

  3. I flagged the three places where a mismatch is most likely, including one where my own implementation diverges from Aleks’s original spec. For the ray-drawing step I deliberately chose not to use the textbook line-drawing method, and instead matched the exact technique the sensors have already used for two months — so the old code and the new code agree with each other. That’s a conscious trade-off, and it needs to be on the record so simulation makes the same choice.

5. Status

  • Training is frozen until simulation answers. They are currently busy with a separate run; I confirmed their current work does not block my question, so there’s no rush on their end.
  • Two control copies of the current model are finishing their runs — for statistics only, not for flight.
  • The full path from here: their answers → a line-by-line cross-check → fixes on my side → retrain with the “smart compass” → a joint verification, number for number, that both worlds match.

The metrics already prove the approach works. The discipline now is to protect that result by making sure the model meets, in the real run, the exact same world it conquered in training.

6. The broader lesson

It is tempting to measure a sprint by lines of code shipped. This entry is the opposite: the highest-value move was a refusal to ship.

The model is fast and it beats the baseline. The weak spot is understood and the fix is known. But a known fix applied against an unverified interface is how you manufacture a silent, expensive bug — exactly the kind we’ve paid for before. So we freeze, we align, and only then do we move. Parity first. Everything else second.

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR