claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

08 — Retraining at the 5000-step limit: more training time did not help

We took our best EXP-7 model, trained on short episodes, and retrained it on long episodes with everything else unchanged. Coverage did not improve.

stablerl-labupdated 2026-05-11T00:00:00.000ZClaudeDroneRLDevLog

What this is: we took our best EXP-7 model, which was trained on short episodes, and retrained it on much longer episodes. Everything else was left identical. Why: when evaluated on long runs, EXP-7 falls short of the classical lawnmower sweep (93.75% vs 98.7% — a 5-point gap). Perhaps EXP-7 never learned what to do beyond its short training horizon, simply because it never saw long episodes during training.

Date: 8 May 2026 Machine: D2


1. What we wanted to test

Context after EXP-7: for the first time, the RL agent had beaten the classical lawnmower heuristic on the short and medium runs (+30 / +17 points). On the longest runs there was still a 5-point gap in favour of the classical method.

EXP-7 was trained on short episodes. In evaluation we let it run roughly five times longer than anything it saw during training. It is reasonable that, once past its training horizon, the agent starts to drift — it had simply never practised such long passes.

The main idea we wanted to check was straightforward: if we retrain the same model on long episodes, will its coverage on the long runs improve by a few points and close the gap to the classical method?

The alternative idea was that the classical lawnmower wins because it can explicitly re-plan when it hits a dead end. If the RL agent gets stuck in loops, more training will not cure that — the fix would have to be architectural rather than just more practice.

2. What we did

Minimal change: we changed only one thing — the episode length cap, lengthening it. Every other setting (the PPO configuration, the feature extractor, the reward) was kept identical to EXP-7. The point was to isolate the effect of a single variable.

Training took about seven minutes on the GPU. Evaluation ran across five unseen maps, five episodes each, at three different run lengths, and took about two minutes.

3. What we saw

Average coverage — nothing changed

Run length EXP-9 (long training) EXP-7 (short training) Δ Lawnmower Gap to lawnmower
Short 55.50% 56.10% −0.6 (noise) 26.1% +29.4 ✅
Medium 84.64% 86.18% −1.5 68.9% +15.7 ✅
Long 93.12% 93.75% −0.6 98.7% −5.6 ❌

The main idea did not hold up. Every difference sat within the noise band, so longer training simply did not help.

By map — only map_03 was slightly better

Map EXP-9 EXP-7 Δ
map_00 93.60% 94.10% −0.50
map_01 94.57% 95.05% −0.48
map_02 92.60% 93.38% −0.78
map_03 (corridors) 92.02% 91.76% +0.26 ⭐
map_04 92.80% 94.48% −1.68

map_03 was the only map with any gain, and even that was inside the noise (σ ≈ 1.3 points).

The main finding — episode lengths

EXP-9 (long training) EXP-7 (short training)
Mean episode length on the long eval 4948 4807

A surprising result: EXP-9 was trained on long episodes, so it should have understood long-horizon behaviour better. Instead the opposite happened:

  • EXP-7 (short training) tends to finish early, reaching 95% coverage and stopping.
  • EXP-9 (long training) tends to run out the full length without reaching 95%.

Interpretation: the long training window made the policy lazier. The model learned to optimise for a slow, full-map traversal rather than for reaching 95% quickly. EXP-7, by contrast, was forced to hurry inside its short episodes, and it carried that brisk behaviour over into the long evaluation runs.

The takeaway is counter-intuitive: more training produced a less effective agent.

4. What this means

This indirectly supports the alternative idea: the gap to the classical method does not close by simply training for longer. Closing it will require architectural or informational changes. The candidates we noted were:

  • A recurrent policy that carries memory of where the agent has already been.
  • A hierarchical controller that issues high-level macro-actions such as “sweep this corridor.”
  • Richer observations that tell the agent where unexplored space remains.
  • A numerical study of the agent’s behaviour to understand exactly what sets EXP-7 apart from EXP-9.

This is now the fifth result in a row, across the “more / longer / different reward” line of experiments, that produced no improvement. All of them point the same way: simple knob-turning will not close the gap; something qualitatively different is needed.

EXP-7 remains the production model.

5. What’s next

  • A numerical analysis: what does EXP-7 do that EXP-9 does not? We plan to look at how often each action is used, how the agent transitions between behaviours, and how effective each action is.
  • A recurrent policy experiment, to give the agent memory.
  • A hierarchical experiment, to introduce macro-actions.
  • An experiment that adds a sense of “where the unexplored space is” to the agent’s observations.

(Later dev-logs cover the results of all these ideas. The short version: they were all flat or worse, right up until the reward-redesign experiment that finally produced the first non-negative result.)

Related files

  • The training configuration for this run.
  • The EXP-9 experiment directory.
  • The EXP-7 experiment directory (used as the baseline for comparison).

Glossary

  • EXP-7 — our strongest baseline, which introduced multi-cell movement.
  • EXP-9 — this experiment, a retrain of EXP-7 on long episodes. It regressed.
  • Episode length cap — an environment setting that forces an episode to end after a fixed number of steps.
  • Rollout length — a PPO setting for how many steps the model collects between weight updates. Not to be confused with the episode length cap.
  • More training → less effectiveness — the surprising pattern seen here, where additional training did not yield a better model at evaluation.
  • Recurrent policy — a neural network that carries memory between steps (such as LSTM or GRU).
  • Hierarchical controller — two-level control: a top level chooses a macro-action and a lower level carries it out.
  • Frontier — the boundary between known and unknown space in a coverage task.
  • Episode length — how many steps the agent took before the episode ended; it can finish before the cap if it reaches 95% coverage.
  • Production model — the model kept as the main baseline against which new experiments are compared.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR