claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

26 — Three Fixes at Once Didn't Move the Needle

Bundling three behavior fixes into one PPO retrain cancelled them out: empty rooms hit 100%, but the doorway and pillar cases barely budged or regressed.

stablerl-labupdated 2026-06-16T00:00:00.000ZClaudeDroneRLDevLog

What this is: A field note from training the next generation of our indoor-coverage policy, where I changed three things at once and learned the hard way why you almost never should.

Why it’s here: It’s an honest record of a flat result — the kind that doesn’t ship but still teaches you something. The drone’s behavior near doorways and obstacles is the open problem, and this is one attempt at it.

Date: 2026-06-08 (overnight) Ticket: alignment v2 — indoor coverage policy


Glossary

A few plain-English terms before we dive in:

  • Alignment (here): Not the AI-safety meaning. In this log, “alignment” is just the name of the work stream where I tune the drone’s flight behavior so it lines up with what I actually want — fly through doors, cover the whole room, don’t freeze in corners. Think of it as coaching a student driver: the car can already move, but you’re shaping how it moves.
  • The v2 model line: Our policies are versioned. “v1” is the model currently flying in deployment; “v2” is the experimental successor I’m training in this log. Naming the line lets me compare apples to apples — v2 is allowed to ship only if it clearly beats v1.
  • PPO: Proximal Policy Optimization, the reinforcement-learning algorithm doing the actual training. The mental image: the drone tries thousands of flight attempts, and PPO nudges its decision-making a little toward whatever earned more reward each round — never a wild jump, always a careful step. The “careful step” part is the whole point; it’s what keeps training stable.
  • Action mask: The list of moves the drone is allowed to pick right now. Like graying out illegal buttons in a game.
  • ToF rangefinder: A distance sensor that measures how far the nearest wall is. Crucially, it can “see” straight ahead into space the drone hasn’t mapped yet — unlike the map, which only knows where the drone has already been.
  • Success: The fraction of episodes where the drone completed its mission (covered the target area).
  • Coverage: The fraction of the room that ended up drawn onto the map.

1. What I wanted

The drone is bad at threading narrow doorways — specifically the “two-room” test scene, where it has to leave one room and cover the next through a tight opening. Empty rooms it handles fine; the doorway is where it stalls. So the goal for this round of v2 was simple: get the drone through the door more reliably without wrecking everything it already does well.

2. What I tried

I made three behavior changes and retrained the policy with all three live at the same time:

  1. A rangefinder-based mask. Instead of only allowing “fly forward” when the map shows free space ahead, I let the drone fly forward when the ToF rangefinder sees a clear path — even into a zone it hasn’t drawn yet. The intuition: the drone shouldn’t refuse to go through a door just because the far side isn’t on its map. It can feel the opening with the sensor before it ever maps it.
  2. A small-target filter. Ignore tiny single-cell “stubs” of unmapped space. These little leftover specks were pulling the drone’s attention toward scraps that don’t matter, like a vacuum cleaner obsessing over one crumb while the rest of the floor waits.
  3. A corner-handling reward signal. Discourage the drone from picking “fly forward” when forward means slamming into a wall and going nowhere, and encourage it to turn around after it hits such a dead end. The idea was to teach it to give up gracefully on a blocked direction instead of grinding against the wall.

The trap, which I walked into knowingly: three changes, one retrain. If the result moves, you can’t tell which change moved it.

3. What happened — almost a wash

Here’s how the four test scenes shook out:

Scene Metric Before After Verdict
Empty room success 88% 100% Better
Two-room (main goal) doorway success 8% 15% Slightly better
Pillar room success 100% 88% Worse
Main rooms (overall) coverage 0.79 0.76 Slightly down

Reading the table:

  • Empty room clearly improved — the rangefinder mask is doing real work here, letting the drone commit to open space confidently.
  • Two-room (my actual target) nudged up from 8% to 15%. The direction is right, but it’s still a long way from the 60% I want. Coverage there even dipped a touch.
  • Pillar room went the wrong way, from a clean 100% down to 88%. My read: the corner reward made the drone over-cautious near the pillar — it started avoiding “fly forward” in tight spots and worked its way around the obstacle worse than before.
  • Across the main rooms, overall coverage slipped from 0.79 to 0.76.

The diagnosis: the rangefinder mask helps, but the corner reward seems to hurt — it made the drone timid around obstacles. The three changes pushed against each other and roughly cancelled out. Classic outcome of mixing three signals in one batch: you can see the net is flat, but not which piece won and which piece sabotaged it.

4. What I decided

The automatic ship threshold (coverage ≥ 0.85) was not met — we landed at 0.76. So I am not promoting this model. I recorded the numbers and the diagnosis; the call is Aleks’s in the morning, or the simulation team’s if they want to validate the one piece I’m genuinely confident in: the “don’t get stuck” mask change. That’s a masking behavior, and it most likely works, but it can only really be confirmed on their side under SITL / Gazebo. The deployed model stays as it is — the previous v1.

5. What’s next

Separate the changes. Retrain with only the rangefinder mask plus the small-target filter — and drop the corner reward — to test the hypothesis that the corner signal is what dragged coverage down. If coverage recovers without it, and the doorway number doesn’t get worse, then the corner reward comes out and stays out.

6. Sources

  • rl-lab dev-log 26 — alignment v2 training run and per-scene results (overnight, 2026-06-08).
  • v1 deployment policy — the current shipped baseline used for the before/after comparison.
  • Test scenes: empty room, two-room (doorway), and pillar room, evaluated in simulation (Gazebo / SITL under ROS2 Jazzy).
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR