claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

06 — Retraining at 3000 steps: the ceiling lies in the action design, not in training

We took the same baseline and retrained it with a longer episode limit to test whether a wider training horizon would raise indoor coverage.

stablerl-labupdated 2026-05-11T00:00:00.000ZClaudeDroneRLDevLog

What this is: we took the same baseline and retrained it, this time allowing much longer episodes during training. Why: several earlier runs all landed in the same low coverage band. One idea was that the training horizon was simply too short. With short episodes the agent learns only a local “head toward the unexplored” habit and never has time to develop a longer-range plan, so we widened the training window to give it that chance.

Date: 8 May 2026 Machine: D2


1. What we wanted

We wanted to check whether the modest coverage at short evaluation lengths came from the agent never seeing long episodes during training. The expectation was that retraining on much longer episodes would teach a longer-range traversal style, lifting coverage at the longer evaluation length and ideally closing the gap toward the classical lawnmower baseline.

2. What we did

The change was deliberately minimal: we extended the per-episode step limit for training and left everything else identical to the baseline. We then ran a full training budget on this longer-episode configuration.

3. What we saw

Step limit This run Baseline (extended eval) Lawnmower Gap to classical
short 18.9% 19.4% 26.1% −7 pp
medium 42.3% ~41% 68.9% −27 pp
long 57.6% ~53% 98.7% −41 pp

The improvement from training on longer episodes was effectively zero once measurement noise is accounted for.

Headline: the gap to the classical baseline actually grows as episodes get longer. The lawnmower heuristic keeps scaling roughly linearly, while the learned policy levels off. That is the opposite of what we had hoped to see.

4. What this means

Across this and the preceding runs, the limiting factor does not appear to live in any of the things we had been adjusting: variations in how the reward is structured, an added exploration incentive, or the length of the training horizon all left the result essentially unchanged.

Instead, the limiting factor seems to be in how actions themselves are defined:

  • A single forward action advances only one cell per step, so even a long episode can only ever touch a small fraction of the open space.
  • On long episodes, once roughly half the map is covered the policy tends to get stuck revisiting areas it has already seen, and its behaviour starts to resemble aimless wandering.
  • There is no action that lets the agent commit to a direction, so it never builds a clean, systematic sweep.

Taken together with the earlier results, this points toward changing the agent’s action design rather than continuing to tune numerical settings.

5. What’s next

The clear priority is a richer forward action that lets the agent move along a direction until it is blocked, rather than crawling one cell at a time. This direction is followed up in the next dev-log entry, and it turned out to be the key. Further variations on the reward and on simply combining a longer horizon with earlier reward tweaks were set aside, since the evidence suggested they would not move the needle.

Related files

  • Training configuration for the long-episode run.
  • Experiment artifacts for this run.

Glossary

  • Episode length limit — the environment setting that caps how many steps a single episode can run.
  • Baseline — our first strong reference policy, against which later runs are compared.
  • Extended eval — evaluating a policy on longer episode limits than it was trained on.
  • Lawnmower — a classical, non-learned coverage heuristic that sweeps the space in regular rows.
  • Sweep pattern — a snake-shaped traversal that covers the area in even rows.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR