claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

14a — Squeezing +5pp From an Existing Model With a Simple Tune

How a small training-knob sweep gave +5pp indoor coverage for free, and why a potential-based reward idea added nothing on top.

stablerl-labupdated 2026-06-16T00:00:00.000ZClaudeDroneRLDevLog

What this is: Eleven days ago we froze a production model that covered 56% of the map in the first 1,000 steps. In about two hours today we found that turning one training knob lifts that to roughly 62% — a free +5pp. We also confirmed that adding a hand-crafted guidance signal to the reward gives nothing on top, which lets us close a whole branch of ideas. Why it’s here: It explains why we chose to “squeeze the current model first, then escalate” instead of jumping straight to a more complex hierarchical setup. This is not academic curiosity — it is free coverage for an autonomous indoor drone in production.

Date: 2026-05-19 Ticket: TASK-RL-SWEEP (find better training settings) + TASK-RL-POTENTIAL (try adding a guidance signal to the reward)


Glossary

  • Training setting (hyperparameter): a dial you set before training starts; the algorithm does not learn it on its own. Think of it like the temperature on an oven — you choose it, the cake does not.
  • Exploration knob: how strongly the algorithm is nudged to keep trying random actions. Turn it up and the drone wanders and explores; turn it down and it commits faster to the strategy it already found.
  • Step-size knob (learning rate): how big a jump the algorithm makes each time it updates the model. Big jumps learn fast but can overshoot a good answer; small jumps are slower but steadier.
  • Coverage: the share of the map the drone has visited by the end of an episode. This is our headline metric.
  • The production model: our model from 2026-05-08 — the first time our learned policy beat the classic lawnmower (zig-zag) baseline.
  • Sweep: systematically trying many combinations of training settings to see which works best.
  • Potential-based guidance: a mathematically safe way to add a hint to the reward. There is a classic 1999 result (Ng et al.) showing this kind of hint, when shaped correctly, can speed learning without changing which behavior is ultimately best.

1. What I wanted to find out

Aleks proposed two cheap checks before we declared the model had hit a ceiling.

Check one — sweep the training settings. Over the previous eleven days, four architectural experiments in a row had come back with no improvement. But all four ran with the same exploration knob and the same step-size knob. Before saying “we’re at the ceiling,” it was worth asking: what if the baseline itself was simply tuned poorly? If some untried combination had a sweet spot, that would be several points of coverage for free. If the surface was flat, the ceiling would be confirmed honestly.

Check two — add a safe guidance signal. This was the one reward-related idea that, by design, cannot break what already works. The hint we tried was, conceptually, “how close is the drone to the nearest cell it hasn’t visited yet?” If that hint carried useful information the network didn’t already have, coverage would rise. If it was redundant, we’d see nothing.


2. What I tried

The sweep

I ran a batch of training runs covering several values of the exploration knob crossed with several values of the step-size knob, each trained for a fixed budget — roughly an hour and a half of GPU time in total. A small orchestrator script kicked them off one after another.

Once a clear winner emerged, I retrained it several more times with different random seeds. That matters: a single lucky run can look great by chance, and repeating it tells you whether the gain is real.

The guidance-signal experiment

I wrote a thin wrapper around the simulation environment that, conceptually, adds the “distance to the nearest unvisited cell” hint to the reward in the mathematically safe form. Then I trained one run using the sweep’s best settings plus this hint. It took about five minutes.


3. What happened

The sweep paid off — about +5pp for free

Coverage at 1,000 steps across the grid of settings (one cell per combination):

            step-size →   small      medium       large
  explore high            54%        61% (best)   60%
  explore mid             59%        61%          55%
  explore (baseline)      54%        57%          54%
  explore very high       45%        46%          44%  ← collapse

The best corner was the lower-exploration, medium-step-size combination. After repeating it with fresh random seeds:

Horizon Mean coverage Change vs production model
@ 1,000 steps ~61.6% +5.5pp
@ 3,000 steps ~91.2% +5.0pp
@ 5,000 steps ~95.0% +1.3pp

The spread across seeds was tiny, so this is a genuine effect, not noise.

A hidden bonus for the real drone: the new model reaches over 95% coverage in roughly 3,700 steps, versus about 4,800 before — about 22% faster. For a battery-powered autonomous drone, faster coverage means less battery spent on the same mission.

Against the classic lawnmower (zig-zag) baseline, we now win clearly early on (+35.5pp at 1k, +22.3pp at 3k). On the long horizon the lawnmower is still slightly ahead at 5k (98.7% vs 95.0%), but we hit the 95% mark about 25% sooner — a better trade on battery.

Why the old setting was holding us back

The production model’s exploration knob was set too high. The algorithm kept being rewarded for trying random actions even after it had already learned an effective zig-zag pattern. Picture a chef who keeps experimenting with a recipe long after it’s already good — the average dish suffers from the random variation. Turning exploration down let the policy lock onto its learned strategy sooner, which lifted coverage.

The very-high exploration corner was a disaster: the exploration nudge drowned out the actual goal, and the drone moved almost at random.

The guidance signal added nothing — as expected

Horizon Sweep winner With guidance signal Difference
@ 1,000 steps 61.6% 60.4% −1.2pp (tiny dip)
@ 3,000 steps 91.2% 91.2% tie
@ 5,000 steps 95.0% 95.0% tie

The reason is informative. Our vision network reads a 64×64 map of where the drone has already been, and a convolutional network over that map is already, implicitly, computing how far the drone is from the nearest unvisited cell. Feeding the same fact in explicitly as a reward hint just duplicates what the network already knows. Early in training that duplication even adds a little noise, which is why we see the small dip at 1k. Over a longer run the safe-shaping property kicks in and the two are equal again, which matches the ties at 3k and 5k.

This is the same lesson as an earlier experiment that fed frontier distance as an observation rather than a reward: the frontier information already lives inside the vision network, so supplying it from the outside is just extra noise.


4. A retrospective note

All four of our earlier “no improvement” results were measured against a baseline that, we now know, was tuned sub-optimally. Each had its own reason for not working — a metric being gamed in one case, redundant memory in another, information already present in the observation in a third. None of that changes their individual conclusions, but the size of the gap to the production model was measured against a weaker version of it. Re-running them on the better baseline isn’t a priority: the squeeze already delivered real production coverage, and moving toward the real drone matters more.


5. What’s next

  • This session: close out the squeeze work, update all the docs, and copy media for the report channel.
  • Next session (Aleks’s direction): sim-to-real domain randomization on the new, better baseline. Concretely, that means training against realistic imperfections so the model survives the jump to hardware:
    • Sensor noise (a distance sensor that’s off by roughly ±10%).
    • Sensor delay (a time-of-flight sensor that reports 50–100 ms late, like real hardware).
    • Servo backlash (turns that land a few degrees off, not perfectly).
    • Goal: a model that works in simulation should work on the real drone without retraining.
  • Reference: arXiv 2409.03930v1 (sim-to-real domain randomization for UAVs).

Sources

  • The sweep orchestrator and its run configs.
  • The thin environment wrapper implementing the safe guidance signal.
  • The summary notes for the sweep ticket (winner details) and the guidance ticket (the no-improvement result).
  • The project plan, where the new production baseline is recorded.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR