What this is: we trained the model on rotated maps only (90°, 180°, 270°, 0°), without mirrors. The question was whether the full map augmentation set lost ground because of the mirrors, and whether rotations alone might do better. Why it matters: a counter-intuitive result — less augmentation turned out worse than more. The lesson: augmentation symmetry matters in itself.
Date: 10 May 2026 Machine: D2
1. What we wanted to check
Context: In an earlier run the full map augmentation set (rotations combined with horizontal and vertical mirrors, giving sixteen map orientations) regressed coverage by a few points. The action analysis turned up something interesting: the baseline model had fixed heavily on a single small counter-clockwise turn, choosing it close to half the time, and the full augmentation set dissolved that fixation back down to roughly a tenth of actions.
That raised the questions we set out to explore here:
- Whether rotations alone (four orientations) — a milder form of augmentation — would keep coverage on par while still breaking the fixation, giving healthy rotation invariance.
- Whether the mirrors had been adding a kind of semantic noise (a corridor opening from the opposite side), so that rotations alone — which are clean isomorphic orientations — might even grow coverage.
- Whether the fixation on the small counter-clockwise turn was purely tied to rotation, in which case it should break with rotations alone.
2. What I tried
The environment had already been set up in a previous session so that rotation and the two mirror directions could be toggled independently. The training script, however, only carried a single combined augmentation switch. I patched it so it would pass the three orientation controls through separately from the config.
I created a rotations-only configuration: map augmentation on, rotations enabled, both mirror directions disabled, with the same exploration settings as the baseline. Everything else was held identical to the baseline using standard PPO settings, training for one million steps across sixteen parallel environments.
Evaluation ran on five fixed maps with five episodes each, sampled at three horizons (1k, 3k, 5k steps), plus an analysis of the action distribution at the longest horizon.
3. What came out
Training: about eight minutes of training, with mean per-episode reward landing below the baseline by roughly six percent. The full augmentation set had trained on par with the baseline — so rotations alone trained worse than the full set.
Coverage (eval):
| Horizon | Rotations only | Baseline | Full map augmentation | Δ vs baseline | Δ vs full set |
|---|---|---|---|---|---|
| 1k | 51.35% | 56.56% | 54.96% | −5.21 pp | −3.61 pp |
| 3k | 76.16% | 86.18% | 83.16% | −10.02 pp | −7.00 pp |
| 5k | 84.34% | 94.12% | 91.04% | −9.78 pp | −6.70 pp |
The key counter-intuitive fact: rotations only (four map variants) gave a worse result than the full set (sixteen variants).
Per map at the longest horizon:
- The open-space map and the corridor map were hurt least, losing about seven points each.
- The remaining three maps were worse, losing roughly nine to fifteen points.
- The worst single map lost close to fifteen points.
Maps with rotational symmetry (the open-space and corridor maps) suffered least — which points to the damage coming specifically from orientational instability.
Action distribution at the longest horizon:
| Action | Rotations only | Baseline | Full map augmentation |
|---|---|---|---|
| Small counter-clockwise turn (the fixation) | 13.4% | 47.7% | 12.0% |
| Small clockwise turn | 39.0% | 16.8% | 37.4% |
| Scan (compensation) | 6.3% | 2.8% | 15.2% |
We confirmed that the fixation on the small counter-clockwise turn broke with rotations alone (down to about a tenth of actions from nearly half at baseline). So the fixation had been tied to rotation — it arose because during training the map always sat in one orientation, which made that single turn the optimal macro-strategy.
But here is the critical contrast: the full augmentation set also broke the fixation, yet the agent compensated for the lost strategy by scanning far more often (more than five times as much) and using more sideways moves. Rotations only broke the fixation too, but gave no compensation — scans barely rose — so the agent was left without a working strategy, and coverage collapsed.
Main takeaway
Augmentation symmetry matters in itself.
| Augmentation type | Variant count | Relative training reward | Coverage drop at 5k |
|---|---|---|---|
| No augmentation (baseline) | 1 | on par | 0 (baseline) |
| Rotations only | 4 | about six percent lower | −9.78 pp |
| Full set (rotations + mirrors) | 16 | on par | −3.08 pp |
Counter-intuitive: less augmentation (four variants) did worse than more (sixteen variants).
A likely mechanism: the full set with mirrors creates geometric fallback routes — mirror invariance lets the agent learn symmetric strategies around the mirror axes. Rotations alone strip those routes away:
- The agent sees four unique orientations and has to learn a rotation-invariant policy.
- Without mirror invariance there are no symmetry axes to compress the representation around.
- The network capacity hits a ceiling and settles into an undertrained, middling policy.
- Without the kind of diversity the full set provides, no compensation strategy emerges — scanning and sideways moves don’t grow.
In the full set, the many orientations and their symmetries let the agent find reliable strategies, leaning much harder on scanning and sideways motion. With rotations alone, four orientations without symmetry leave it undertrained.
Methodological lesson
The pattern from earlier runs continues: literature claims transfer only when the applicability conditions match, and when an input-importance test isolates the right axis. The rotations-only run was designed to separate rotation invariance from any semantic noise the mirrors might add; the result showed that augmentation symmetry itself matters, and that not every subset works.
For production
- Don’t do partial augmentation — either the full set of orientations or none.
- Rotations-only as an additive component is discarded.
- Future augmentation experiments should explicitly test geometric coverage symmetry rather than the intuition that “less is safer.”
4. Sources
- arxiv 1812.02900 — Cobbe et al. 2018, the augmentation idea in RL.
- arxiv 2306.16978 — Niroui CPP RL, the basis for our exploration reward.
- Internal notes: the earlier full map-augmentation run for a direct comparison, and the earlier run that established the same “literature doesn’t transfer without checking” pattern.
5. What’s next
- An entropy sweep aimed directly at the fixation, testing the idea that the issue is agent confidence rather than orientation.
- A network scale-up to add capacity, which might lift the undertraining.
- A frontier multi-cell idea — harder, but with high potential.
- Longer training, kept low priority since an earlier run showed longer training made the agent lazier.
Glossary
- Augmentation: artificially expanding the training set through transforms such as rotations and mirrors.
- Cobbe et al. 2018: the paper showing augmentation as the main cure for the generalization gap in RL.
- Generalization gap: the quality gap between training maps and eval maps. If a model has memorized its training maps, the gap is large.
- Exploration reward (Niroui-style): extra reward for newly visited cells, which encourages exploring the map.
- Fixation: when a model picks one action unnaturally often, even when it isn’t optimal on the current map.
- Scan: active sensor scanning, gathering data without moving.
- Strafe: sideways motion perpendicular to the heading.
- Mirror invariance: the property of working equally well on a map and its mirror image.
- Per-map analysis: breaking results down per eval map rather than only averaging.