Last week we wired up the simulation. This week we used it. By Friday, the first PPO checkpoint had been trained, evaluated, and visibly underperformed the lawnmower baseline — which was the expected result and a good sign, because if first-attempt PPO had beaten a hand-tuned classical planner, something would have been suspiciously wrong with our evaluation harness.
What landed
- First PPO training run completed end-to-end on the
warehouse_v2map. 200k training steps, MlpPolicy,γ=0.99, defaultentropy=0.01, no curriculum. Coverage ~0.62 vs the classical baseline’s ~0.81. Full writeup: 02-first-ppo-run. - Frontier-reward variant. First non-trivial reward design — a bonus for moving toward unexplored frontier cells rather than just coverage credit on the cell currently under the drone. Coverage jumped to ~0.74. The lesson: greedy coverage rewards encourage local sweeping; frontier rewards encourage exploration. Both have a place. Details: 03-frontier-reward.
- RND + rllte experiment. Random Network Distillation as an intrinsic-reward bonus, implemented via rllte. The hypothesis was that RND would unlock harder maps where extrinsic reward signal is sparse. Initial results were noisy; we shelved it for now and came back to it later in the experiment sequence. Details: 04-rnd-rllte.
The bigger picture
This week answered a question that mattered: can our pipeline produce a policy that learns anything? Answer: yes. The PPO checkpoint isn’t competitive with classical baselines yet (it won’t be until EXP-7 several weeks later — see the RL vs lawnmower benchmark), but it’s clearly learning. Coverage curves climb, value loss decreases, episode length stabilizes — all the basic signs that the loop is closed and the agent is actually optimizing.
What it doesn’t yet do: handle out-of-distribution maps gracefully (that’s a recurring theme through the whole experiment sequence), survive sensor noise (the σ-sweep noise-robustness tests come in dev-log entry 21), or beat the classical baseline (EXP-7 closes that gap).
A note on baselines
We compare every RL run against a hand-tuned classical baseline (lawnmower). This is the single most useful methodological decision we made early. Without it, every RL result looks meaningful in isolation; with it, you immediately see whether “0.74 coverage” is good (better than baseline) or embarrassing (worse than baseline). For most weeks this month, “embarrassing” was correct — and that’s fine, because the baseline pins a ceiling worth aiming at.
Where this leads
- Week 3 — Gazebo physics tuning
- PPO hyperparameters — what the first runs used and how we tuned from there
- Reward functions overview — the design space we’re exploring
- The full RL dev-log index — every experiment, in order, with results