claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

15 — Growing the map bank from 64 to 256: is the limit in the data or the method?

We took our best baseline and retrained it on 256 random maps instead of 64. A single change — the size of the training set.

stablerl-labupdated 2026-05-11T00:00:00.000ZClaudeDroneRLDevLog

What this is: we took our best baseline and retrained it on 256 random maps instead of 64. A single change — the size of the training set. Why: five experiments in a row showed the same pattern: “trains great, evaluates poorly.” That is the symptom of either “too few maps” or “the method cannot generalize.” This experiment was meant to tell us which.

Date: 9 May 2026 Machine: D2


1. What we wanted to check

Each of the last five experiments showed the same pattern: the model earned more and more reward during training, but on new maps the result did not grow (or even fell). There were two ways to read this:

  • Interpretation A — too few maps in training. The model “memorizes” 64 maps and never learns to generalize to new ones. The remedy would be more maps.
  • Interpretation B — the method has topped out. The overall design simply cannot do better, no matter how much data we feed it. The remedy would be a different structure (a layered approach, ordered training).

Cost of the experiment: about 30 minutes. Cost of guessing wrong: a month spent on the wrong strategy.

2. What we did

We generated 256 maps with the same procedure as before (five minutes): same type, same wall density, same generator. Simply four times as many.

We trained the exact same setup as our best baseline, with the same settings and the same budget. The only thing that changed was the input — the maps folder went from 64 to 256.

During training: reward rose to 1990 versus 1949 for the baseline, about +2%. Slightly better — the model handled the larger set without trouble.

We evaluated on the same five unseen maps used for the baseline, so the comparison stayed fair.

3. What we saw

Average coverage — it got worse

Step limit 256 maps (new) 64 maps (baseline) Difference
1000 54.75% 56.10% −1.35 points
3000 82.57% 86.18% −3.61 points
5000 90.66% 93.75% −3.09 points

So the “too few maps” reading did not hold up: four times more maps, and the result did not rise — it slipped a little. That points instead toward the method having reached its limit: somewhere in the overall design there is a barrier that more data does not break.

Per map — the hardest maps suffered most

Map 256 maps 64 maps Difference
map_00 89.95% 94.10% −4.15
map_01 (easy) 95.06% 95.05% +0.01 (parity)
map_02 (hard) 86.67% 93.38% −6.71 ⚠
map_03 (corridors) 88.66% 91.76% −3.10
map_04 92.96% 94.48% −1.52

A sign of life: on map_01 the model sometimes finishes coverage before the step limit (mean episode length 4864 versus always 5000 in the prior variant). So training did learn something useful — it is just slower.

A pattern that has now become systemic

Experiment Training reward Coverage on new maps
Baseline (64 maps) 1949 56.10%
Recurrent memory parity 54.62% (−1.48)
Extra observation cue 2030 (+4%) 54.67% (−1.43)
Fewer actions 2070 (+6%) 54.80% (−1.30)
256 maps 1990 (+2%) 54.75% (−1.35)

Six rounds in a row: training holds parity or improves, while evaluation lands one to three points lower. Whatever we add, this limit does not move.

4. What this means

The strategic takeaway: the “ordinary” machine-learning levers are exhausted:

  • ❌ Memory (recurrent network)
  • ❌ Extra information in the observations
  • ❌ Restricting the set of available actions
  • ❌ Fewer actions
  • ❌ More data (64 → 256 maps)

What is left:

  • A layered control structure — a qualitatively different design where a high level decides what to do and a low level decides how. Not “more,” but “different.”
  • Ordered training — teaching the model from easy to hard, not with more maps but in the right order.
  • Bridging from simulation to reality — a different axis altogether, not about coverage.

A new working rule: before attempting a layered approach, do a deliberate review of 10–15 sources first (rule introduced 9 May). After six dead ends, 45 minutes of reading the literature saves hours of failed code.

5. What’s next

  • The layered approach is the top priority, with no real alternative. Literature review first.
  • Ordered training moves up from low to medium priority as a fallback.
  • Optional: check whether map_02 is simply undertrained — at −6.71 it is the worst case. If retraining with double the budget catches it up, we just ran out of time. If not, it confirms the design limit. About 12 minutes, kept in the backlog.

Related files

  • Training maps for this run (256 maps)
  • Experiment artifacts for this run
  • Evaluation script (extended to save results for different step limits into separate files)

Glossary

  • DR (Domain Randomization) — the practice of training across many different maps so the model learns to generalize rather than memorize.
  • Baseline — our current best-performing model.
  • Training reward vs evaluation coverage — two different metrics. The first shows “how training is going,” the second “does the model work on new data.”
  • Train-up / eval-flat pattern — training keeps improving while evaluation does not (or worsens). A classic symptom of overfitting or memorization.
  • The limit — the level above which results will not rise without a qualitative change of approach.
  • Layered control — two levels of decision-making: the top picks a strategy (for example, “head to the north-east corner”), the bottom carries it out (turns, steps).
  • Ordered training — teaching “easy to hard,” the way school works.
  • Literature review — reading 10–15 sources before starting an experiment.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR