What this is: we took the same baseline and retrained it, this time allowing much longer episodes during training. Why: several earlier runs all landed in the same low coverage band. One idea was that the training horizon was simply too short. With short episodes the agent learns only a local “head toward the unexplored” habit and never has time to develop a longer-range plan, so we widened the training window to give it that chance.
Date: 8 May 2026 Machine: D2
1. What we wanted
We wanted to check whether the modest coverage at short evaluation lengths came from the agent never seeing long episodes during training. The expectation was that retraining on much longer episodes would teach a longer-range traversal style, lifting coverage at the longer evaluation length and ideally closing the gap toward the classical lawnmower baseline.
2. What we did
The change was deliberately minimal: we extended the per-episode step limit for training and left everything else identical to the baseline. We then ran a full training budget on this longer-episode configuration.
3. What we saw
| Step limit | This run | Baseline (extended eval) | Lawnmower | Gap to classical |
|---|---|---|---|---|
| short | 18.9% | 19.4% | 26.1% | −7 pp |
| medium | 42.3% | ~41% | 68.9% | −27 pp |
| long | 57.6% | ~53% | 98.7% | −41 pp |
The improvement from training on longer episodes was effectively zero once measurement noise is accounted for.
Headline: the gap to the classical baseline actually grows as episodes get longer. The lawnmower heuristic keeps scaling roughly linearly, while the learned policy levels off. That is the opposite of what we had hoped to see.
4. What this means
Across this and the preceding runs, the limiting factor does not appear to live in any of the things we had been adjusting: variations in how the reward is structured, an added exploration incentive, or the length of the training horizon all left the result essentially unchanged.
Instead, the limiting factor seems to be in how actions themselves are defined:
- A single forward action advances only one cell per step, so even a long episode can only ever touch a small fraction of the open space.
- On long episodes, once roughly half the map is covered the policy tends to get stuck revisiting areas it has already seen, and its behaviour starts to resemble aimless wandering.
- There is no action that lets the agent commit to a direction, so it never builds a clean, systematic sweep.
Taken together with the earlier results, this points toward changing the agent’s action design rather than continuing to tune numerical settings.
5. What’s next
The clear priority is a richer forward action that lets the agent move along a direction until it is blocked, rather than crawling one cell at a time. This direction is followed up in the next dev-log entry, and it turned out to be the key. Further variations on the reward and on simply combining a longer horizon with earlier reward tweaks were set aside, since the evidence suggested they would not move the needle.
Related files
- Training configuration for the long-episode run.
- Experiment artifacts for this run.
Glossary
- Episode length limit — the environment setting that caps how many steps a single episode can run.
- Baseline — our first strong reference policy, against which later runs are compared.
- Extended eval — evaluating a policy on longer episode limits than it was trained on.
- Lawnmower — a classical, non-learned coverage heuristic that sweeps the space in regular rows.
- Sweep pattern — a snake-shaped traversal that covers the area in even rows.