What this is: we added a new move to the agent’s repertoire — “go forward until you hit a wall.” A single command can now travel dozens of cells, where previously the only forward move advanced one cell at a time. Why: four runs in a row, each tuning incentives, exploration, and episode length differently, could not push indoor coverage above roughly 19%. A simple physical estimate showed that with one-cell motion it is impossible to cover a 64×64 map within a 1000-step budget. That points to a design limit, not a learning failure.
Date: 8 May 2026 Machine: D2
1. What we wanted to check
Context after four flat runs:
| Run | Coverage @1000 steps | Gap to classical |
|---|---|---|
| Frontier + CNN + randomization | 19.4% | −6.7 pp |
| Frontier inverse-distance | 19.0% | −7.1 |
| Curiosity-driven exploration | 17.8% | −8.3 |
| Longer-budget retrain | 18.9% | −7.2 |
| Lawnmower (classical) | 26.1% | — |
The learned policy trailed the classical reference by roughly 1.4–1.7× across every horizon. No combination of incentives and time budget closed the gap.
The estimate: a 64×64 map is 4096 cells (walls plus open space), of which about 3800 are free. The drone advances one cell per tick, so within 1000 ticks it simply cannot visit more than 1000 cells — and fewer once turns and scans are counted.
Working idea: the problem lay not in learning but in how moves were defined. If a single command could carry the agent forward until it meets an obstacle, one tick might cover dozens of cells, removing the structural ceiling.
2. What we did
Environment change: we introduced a new move that advances while the path ahead is clear, crediting each newly visited cell exactly as a normal step would, and stopping with the usual collision penalty when it meets a wall. A safety cap limits each invocation to at most 64 cells.
Everything else stayed the same: the same PPO setup, the same feature extractor, and the same randomization across 64 maps as in the previous runs. The only change was widening the discrete set of moves by one and adding the corresponding handling code.
Training ran for one million steps across 16 parallel environments, finishing in about seven minutes.
3. What we saw
Coverage — past the classical baseline
| Step limit | This run | Previous best | Lawnmower | RL / lawn |
|---|---|---|---|---|
| 1000 | 56.10% | 18.9% | 26.1% | 2.15× ⭐⭐ |
| 3000 | 86.18% | 42.3% | 68.9% | 1.25× ⭐ |
| 5000 | 93.75% | 57.6% | 98.7% | 0.95× |
Highlights:
- At 1000 steps: +30 pp over the classical reference (56% vs 26%), about 2.15× the coverage.
- At 3000 steps: +17 pp over the reference (86% vs 69%), about 1.25×.
- At 5000 steps: still about 5 pp behind the reference, though episodes often end early because the model reaches roughly 95% and stops.
Against the previous best run, this is close to 3× the coverage at 1000 steps — purely from giving the agent a richer set of moves.
By map
- Open map — the leader at 60.6% / 90.5% / 95.0%. Long corridors let the forward move pay off the most.
- Narrow-corridor map — the hardest at 52.4% / 81.7% / 91.8%. The forward move often meets a wall quickly and stops short.
Training dynamics
- Per-episode return climbed from the low tens at the start to a stable band by the end of training, without signs of degeneration.
- Episode length stayed at the full budget on the short episodes.
- The policy grew steadily more decisive over training while still retaining variety in its choices.
Eval video
In the evaluation clip (TASK-RL-EXP-7_eval.gif) the agent shows a learned snake-like traversal: long straight passes, a turn, then a fresh pass. The policy worked out the classical sweeping pattern on its own through training, with no explicit programming of that behaviour.
4. What this means
Headline: the limiting factor was the move set, not the learning. With the design ceiling removed, there is room to keep improving.
How it lined up with our expectations:
- The expectation of a roughly threefold gain from the new move held up.
- The expectation that the new move would dominate the agent’s behaviour held up indirectly; a precise count was left for the following experiment.
- On calibration: at the 3000-step budget we expected to merely match the reference and instead beat it; at the 5000-step budget we expected to come close to its near-complete coverage and fell about 5 pp short.
Against the literature:
- A published study on RL coverage path planning with Gazebo on a 50×50 grid (arXiv 2110.09018) built a multi-cell move into the design from the outset and reached roughly 85% coverage in a reasonable number of steps. Our result is comparable.
- The lawnmower heuristic implicitly relies on the same trick: a breadth-first traversal is forward-until-collision plus replanning at dead ends. Our policy learned this through training rather than having it hard-coded.
On the goal of beating the baselines by a small margin — we exceeded it by a wide margin on the short budgets, while the longest budget still leaves a gap to close.
5. What’s next
- A numerical look at how often the new move is chosen and how much it contributes to coverage. (As it turned out, it was used a modest fraction of the time yet drove almost all of the coverage — see dev-log/09.)
- A retrain at the 5000-step budget, in case the policy simply never learned what to do beyond the first 1000 steps. (As it turned out, longer-horizon training did not help — see dev-log/08.)
- Hierarchical control — the forward move is already a natural building block for a “pick a region, then sweep it” hierarchy, a candidate for future work.
Related files
- The 2D drone environment (the new move).
- The PPO training configuration for the multi-cell setup.
- The experiment directory for this run.
~/drone_media/rl/TASK-RL-EXP-7_eval.gif— traversal video.
Glossary
- Action space — the set of moves an agent can choose from. Here it grew by one to include the new forward move.
- Multi-cell forward — the “forward until obstacle” move; one command corresponds to many steps in the environment.
- Forward-until-collision — the precise name: advance until you meet something.
- Lawnmower — a classical, non-RL coverage heuristic with a lawnmower-shaped traversal and replanning at dead ends. Coverage of 26.1% / 68.9% / 98.7% at the 1k/3k/5k budgets — the reference we aim to match.
- Boustrophedon (literally “as one ploughs with an ox”) — a snake-shaped traversal pattern. The lawnmower heuristic uses it, and our policy learned it on its own.
- Domain randomization — training across many randomized maps so the model generalizes.
- Feature extractor — the network component that fuses distance readings, servo angle, and the visited-cell grid into shared features.
- PPO — Proximal Policy Optimization, the reinforcement-learning algorithm used here.
- Entropy — a measure of how decisive the policy is; high entropy means varied choices, low means deterministic ones.
- Stochastic eval — evaluation that samples moves by their probabilities rather than always taking the most likely one.