claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

May 2026, Week 2 — First RL runs

The first PPO training run, frontier-reward variant, and an RND-based exploration experiment — the RL framework starts producing.

stubdoc-seoupdated 2026-05-11T00:00:00.000ZClaudeDroneDevLogRL

Last week we wired up the simulation. This week we used it. By Friday, the first PPO checkpoint had been trained, evaluated, and visibly underperformed the lawnmower baseline — which was the expected result and a good sign, because if first-attempt PPO had beaten a hand-tuned classical planner, something would have been suspiciously wrong with our evaluation harness.

What landed

  • First PPO training run completed end-to-end on the warehouse_v2 map. 200k training steps, MlpPolicy, γ=0.99, default entropy=0.01, no curriculum. Coverage ~0.62 vs the classical baseline’s ~0.81. Full writeup: 02-first-ppo-run.
  • Frontier-reward variant. First non-trivial reward design — a bonus for moving toward unexplored frontier cells rather than just coverage credit on the cell currently under the drone. Coverage jumped to ~0.74. The lesson: greedy coverage rewards encourage local sweeping; frontier rewards encourage exploration. Both have a place. Details: 03-frontier-reward.
  • RND + rllte experiment. Random Network Distillation as an intrinsic-reward bonus, implemented via rllte. The hypothesis was that RND would unlock harder maps where extrinsic reward signal is sparse. Initial results were noisy; we shelved it for now and came back to it later in the experiment sequence. Details: 04-rnd-rllte.

The bigger picture

This week answered a question that mattered: can our pipeline produce a policy that learns anything? Answer: yes. The PPO checkpoint isn’t competitive with classical baselines yet (it won’t be until EXP-7 several weeks later — see the RL vs lawnmower benchmark), but it’s clearly learning. Coverage curves climb, value loss decreases, episode length stabilizes — all the basic signs that the loop is closed and the agent is actually optimizing.

What it doesn’t yet do: handle out-of-distribution maps gracefully (that’s a recurring theme through the whole experiment sequence), survive sensor noise (the σ-sweep noise-robustness tests come in dev-log entry 21), or beat the classical baseline (EXP-7 closes that gap).

A note on baselines

We compare every RL run against a hand-tuned classical baseline (lawnmower). This is the single most useful methodological decision we made early. Without it, every RL result looks meaningful in isolation; with it, you immediately see whether “0.74 coverage” is good (better than baseline) or embarrassing (worse than baseline). For most weeks this month, “embarrassing” was correct — and that’s fine, because the baseline pins a ceiling worth aiming at.

Where this leads

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR