This page gives a high-level overview of the reinforcement-learning work behind ClaudeDrone’s autonomous 2D indoor mapping. Each stage of the effort has been published in full as its own article in the RL dev-log. The work follows a familiar research arc: stand up the training infrastructure, get a first policy learning, then iterate on the reward design, the agent’s inputs, and robustness, validating each idea against measured outcomes before keeping or discarding it.
Throughout, the policy is trained with PPO and evaluated in a Gazebo-based simulator targeting deployment on ROS2 Jazzy. Progress is judged against a classical coverage baseline and against the agent’s own coverage and efficiency metrics rather than against any single internal target.
Chronology
The sequence below summarises each published entry. Early entries establish the baseline; the middle entries explore reward and input changes; the later entries focus on robustness and bringing the policy closer to real hardware.
| # | Stage | Article |
|---|---|---|
| 01 | Bootstrap: standing up the training infrastructure from scratch | 01-bootstrap |
| 02 | First PPO run | 02-first-ppo-run |
| 03 | Frontier-based reward | 03-frontier-reward |
| 04 | Curiosity-driven exploration via RND | 04-rnd-rllte |
| 05 | Lawnmower coverage baseline | 05-lawnmower-baseline |
| 06 | Retrain with a longer episode budget | 06-max-steps-3000-retrain |
| 07 ⭐ | Multi-cell forward action — a notable step; this pilot beat the classical baseline by ×2.15 | 07-multi-cell-forward-action |
| 08 | Retrain with an extended episode budget | 08-max-steps-5000-retrain |
| 09 | Action-distribution analysis | 09-action-distribution-analysis |
| 10 | Soft action masking — no measurable gain | 10-action-masking-soft-null |
| 11 | Hard-mask reward gaming — no measurable gain | 11-hard-mask-gaming-null |
| 12 | Recurrent policy — no measurable gain | 12-recurrent-policy-null |
| 13 | Frontier-observation regression | 13-frontier-obs-regression |
| 14 | No-scan input-importance test | dev-log 14 |
| 15 | Extended domain randomization — the data-ceiling idea did not hold up | 15-dr-extended-data-ceiling-disconfirmed |
| 16 | N-step advantage estimation — the expected pattern did not hold up | 16-nstep-gae-saw-disconfirmed |
| 17 | Total-variation reward — partially positive result | 17-tv-reward-partial-positive |
| 18 | Total-variation reward sweep — a balanced setting was selected | 18-tv-reward-lambda-sweep |
| 19 | Sensor input-importance test (TF-Luna + servo can be dropped) | dev-log 19 |
| 20 | Map-augmentation removed the reward mechanism — no measurable gain | 20-map-aug-tv-mechanism-destroyed-null |
| 21 ⚠ | Sim-to-real noise sweep — one sensor channel proved critical (−79 pp) | 21-sim-to-real-noise-sweep-vl53_0-critical |
| 22 | Rotation-only augmentation — deep regression, confirming that symmetry matters | 22-rotate-only-aug-deep-regression-augmentation-symmetry-matters |
| 23 | Session-crash post-mortem and fix | 23-session-crash-post-mortem-setsid-fix |
| 24 | Entropy-coefficient curve — mapping the limits of further tuning | 24-ent-coef-curve-closed-plateau-cliff |
Outcomes at a glance
The headline result was the multi-cell forward action in stage 07, which more than doubled coverage efficiency relative to the classical baseline. Later robustness work surfaced an important finding: one sensor channel turned out to be a limiting factor under realistic noise, with a −79 pp drop when it was degraded. Several promising ideas were tested and set aside when they showed no measurable benefit, which is a normal and valuable part of the iteration.
Media
The work is illustrated by 26 evaluation animations and graphs captured during training and evaluation runs.