What this is: we replaced the regular feedforward network with a recurrent one (an LSTM). A recurrent network can remember past observations and carry context across steps. It is a standard approach in reinforcement learning for partially-observable tasks. Why: in our prior analysis it looked like the drone was missing a sense of memory of where it had already been. A recurrent policy seemed like the natural way to address that.
Date: 8 May 2026 Machine: D2
1. What we wanted to test
After closing out the earlier action-shaping work, we turned to changes in the network architecture itself.
A recurrent (LSTM) policy is a standard tool here:
- It is not prone to the network gaming the objective, because it adds no explicit restriction on behaviour.
- Its hidden state can, in principle, hold a sense of “where I have been” more flexibly than any explicit memory we hand it.
- Earlier analysis showed that the drone’s long-forward moves covered only a handful of cells on average, far short of what was geometrically possible. Our hunch was that the drone “forgets” that the cell directly ahead of it is a wall.
What we expected:
- Best case: coverage rises into the high-nineties, closing the gap to the classical reference heuristic.
- Worst case: the recurrent network simply duplicates memory the drone already has, and the result is null.
2. What we did
We trained a recurrent variant of our PPO policy, reusing the same observation encoder as the baseline. A short smoke test confirmed the setup ran, and we then trained for a full run of one million steps.
Training took substantially longer than the feedforward baseline — roughly four and a half times slower per run. The recurrent policy also reached a modestly higher training score than the baseline, suggesting it fit the training distribution somewhat better.
3. What we saw
Average coverage — parity within noise
| Step limit | Recurrent | Baseline EXP-7 | Δ |
|---|---|---|---|
| 1000 | 54.62% | 56.10% | −1.48 |
| 3000 | 84.93% | 86.18% | −1.25 |
| 5000 | 93.42% | 93.75% | −0.33 |
All differences sit within measurement noise (roughly one point of standard deviation).
Gap to the classical lawnmower reference: about −5.3 points at 5000 steps — the same as the baseline. The recurrent policy did not move it.
One odd outlier
On one map at the short limit (1000 steps), the recurrent policy scored 40.07% against 53.48% for the baseline — a −13.4 point drop. This may be a case where the recurrent network’s initial state correlates poorly with that specific map’s geometry (a narrow corridor, perhaps). At 5000 steps the outlier disappears.
4. What this means
The headline: the drone’s observations already contain explicit memory — a fine-grained map marking which cells have and have not been visited. That is a large, precise memory store the policy can read directly.
A recurrent hidden state is, in effect, a compressed representation of that same information. It is not new information — it is a compressed duplicate.
The drone’s long-forward move is itself a form of temporal abstraction: a single chosen action unrolls into many steps in the environment. The drone does not need a recurrent memory to remember “one step ago” — a single action already carries it across many cells.
What our earlier analysis already showed: the drone had already learned a snake-shaped sweep of the space without any recurrent memory. Memory was not the limiting factor.
The real limiting factor (a diagnosis for future experiments): by around 5000 steps the drone runs out of obvious places to go — much of the map is already covered and many routes are exhausted. This is an informational problem (where is the unexplored space?) or a structural one (how to switch between separate unvisited regions), rather than a problem of memory.
Compared with the literature: published work has reported gains from recurrent policies on coverage tasks. But in those settings the observations did not include explicit memory. Ours do, so that reported benefit does not transfer to our case.
5. What’s next
Three architectural changes in a row all landed at parity or mild regression, so the prior baseline remains our production model.
We are pivoting away from architecture and toward the information the drone receives — giving it a clearer signal of where the unexplored space lies, and exploring a more structured, hierarchical approach to navigation.
Related files
- Recurrent training configuration and scripts
- The associated experiment directory
Glossary
- LSTM (Long Short-Term Memory) — a type of recurrent neural network that can hold context from past steps. A standard choice for sequential tasks.
- Recurrent network — a network with memory, via feedback connections.
- Feedforward — a network without memory; the regular kind.
- Hidden state — the internal “memory” of a recurrent network, carried between steps.
- sb3-contrib — an extension of Stable-Baselines3 (our reinforcement-learning library) with more advanced algorithms.
- Visited grid — the map the drone keeps of which cells it has already covered.
- Temporal abstraction — a single action that unrolls into many environment steps, such as a forward move that continues until it reaches an obstacle.
- Boustrophedon — a back-and-forth, snake-shaped traversal.
- Lawnmower — a classical coverage heuristic we use as a reference.