claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

23a — Rewarding Progress: How Step-by-Step Praise Cracked the Doorway Room

An indoor coverage policy hits 10/10 on the two-chamber room and beats the classic planner once we rewarded real progress toward the unknown.

stablerl-labupdated 2026-06-16T00:00:00.000ZClaudeDroneRLDevLog

What this is: A field report from our reinforcement-learning lab about a single design change that finally fixed the hardest indoor room — and, almost as a side effect, pushed the policy past the classic planner on speed. Why it’s here: It’s a clean, repeatable lesson about when extra encouragement helps a learning agent and when it just adds noise. That distinction matters far beyond drones.

Date: 2026-06-07 Ticket: RL coverage policy — indoor exploration


Glossary

Before the story, a few plain-English terms. No math required.

  • Coverage policy. The “brain” that decides where the drone flies next so it sees as much of an unknown indoor space as possible, as fast as possible. Think of a person walking into a dark house and trying to peek into every room with the fewest steps.
  • The frontier (the “unknown edge”). The boundary between the part of the map the drone has already seen and the part it hasn’t. Good explorers are drawn toward this edge — that’s where new information lives.
  • The two-chamber room. Our nemesis test: two open areas joined by a single narrow doorway. The trap is obvious to a human (“walk through the door”) and surprisingly hard for a learning agent, which loves to spin in a corner of the first chamber instead of committing to the gap.
  • The classic planner. A traditional, hand-coded exploration algorithm. It’s our yardstick: if the learned policy can’t beat it, the learning wasn’t worth it.
  • The enfilade. A chain of rooms strung one after another, where the only way to reveal the last room is to pass through every room before it. The toughest floor plan for any explorer.
  • Snapshot-before-the-move. When we measure how much “closer to the unknown” a move got the drone, we measure against the map as it was before the drone acted — not after. This tiny detail turns out to be the whole anti-cheating story (more below).

1. The change, in one sentence

For a long time we tried to tell the drone where the interesting unexplored space was — a kind of directional hint baked into how it scored situations. It didn’t take. The room with the doorway stayed broken.

So we flipped the framing. Instead of describing the road, we paid the agent a small bonus every time it actually shortened the distance to the unknown edge. Move that gets you genuinely closer to undiscovered space? Good, here’s a little credit. Move that wanders sideways or back into already-seen rooms? Nothing.

That’s the entire idea, stated in plain prose. We rewarded real forward progress toward new information, and we discouraged dithering. We deliberately avoid spelling out the exact internal recipe here — what matters is the shape of the incentive, not the bookkeeping behind it.


2. The doorway room, finally solved

This was the target case, and it went from “sometimes” to “always.”

Two-chamber room Result
First attempt — directional hint baked in 8 of 10, occasionally got stuck
Second attempt — a subtler “whisper” of where to go 0 of 10, spun in a corner
Latest — bonus for real progress 10 of 10, in 91 steps (classic planner: ~325)

The most striking part: there were two episodes where both earlier versions spun in place for the full thousand-step budget and never escaped. The new policy closes those same two rooms in 134 and 123 steps.

Comparison clips live next to the old ones: ab_v15c_multiroom00.gif and ab_v15c_two_chambers.gif — the side-by-side really sells it, the old runs circling endlessly while the new one walks straight through the gap.


3. The wider scoreboard

The fix wasn’t a one-room fluke. Across all our test arenas:

  • Simple rooms: 0.918 of the map covered, 30 of 30 successful runs — the first perfect sweep of the whole sprint.
  • Apartment layouts: 43 of 50 successful (up from 31 and 12 in the two earlier versions). On speed, it reaches 80% of the map in about 284 steps on average — and that average is pessimistic, because we counted every failure as the maximum thousand steps. The classic planner needs roughly 450. In other words, the policy is now honestly faster than the classic planner, with no asterisk about “when it happens to work.”
  • Still on the to-do list: 7 failures out of 50, all of them long enfilades — those daisy-chained rooms. That’s the next thing to chase, not a hole in the idea.

4. The anti-cheating guarantee (built in, and tested)

There’s an obvious way for a clever agent to game a “progress bonus”: just park near the entrance and collect credit every time the map ticks over on its own, without actually going anywhere. We worried about exactly this from the start.

The defense is the snapshot-before-the-move rule. Progress is always scored against the map as it looked before the drone acted. Stand still and the measured progress is zero, no matter what’s happening to the map around you. Loitering earns nothing.

We didn’t just trust that argument — we wrote a test for it. It compares, down to the last fraction, what the agent earns in two parallel worlds (one with the bonus, one without) on pure in-place rotations. The numbers match exactly, which means the bonus genuinely contributes nothing when the drone isn’t making real headway.


5. The lesson worth keeping

Here’s the part that travels beyond this project.

Back in May, a very similar “reward the progress” idea failed. Same instinct, no payoff. Why the opposite result now?

Because last time the agent already saw that signal with its own eyes — the information was right there in what it perceived, so the extra bonus just duplicated something it had. It was noise on top of a clear channel.

This time the signal was scarce. The agent couldn’t easily perceive how close it was getting to the unknown, so the bonus filled a real gap rather than echoing an existing one.

The rule, in one line: this kind of encouragement cures a hunger, it doesn’t pile extra on a full plate. Before you add an incentive to a learning agent, ask whether the information is missing or merely under-weighted. If it’s already visible, you’re adding clutter. If it’s genuinely absent, you might be handing the agent the one thing it needed.


6. Where this goes next

The remaining failures are all enfilades — those long chains of rooms. The progress signal points toward the nearest slice of unknown, which is exactly right in open layouts but can leave the agent satisfied too early when the real prize is three rooms deeper. That’s the next puzzle on the bench. The doorway, at least, is closed for good.

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR