What this is: we added map rotations and mirrors to training — each map could appear in any of its rotated and mirrored orientations. The idea: the model should learn to work on any orientation of a map rather than a memorized shape. Why: in an earlier experiment a reward-shaping variant gave a small coverage gain. We wanted to see whether map augmentation could amplify that result, since augmentation is a standard generalization trick in the literature.
Date: 10 May 2026 Machine: D2
1. What we wanted to test
Context: in the previous experiment our reward-shaping variant produced a small coverage gain over the earlier production baseline. Looking at how the model behaved, we noticed something odd: it leaned very heavily on one particular turn direction. That could be either:
- a genuinely smart traversal strategy that the reward had taught, or
- overfitting to the specific orientations of the training maps.
Hypothesis: if we rotate and mirror the maps during training, the model can no longer lean on a specific orientation. So:
- If the helpful behavior is real and transferable, coverage should hold steady or improve.
- If it was overfitting, coverage should fall — but we would at least know we had an honest, orientation-invariant model.
We went in expecting augmentation to help, based on prior generalization work, and treated the result as a genuine test of whether that habit was a real mechanism or just memorization.
2. What we did
We built a new training environment that, on every reset, randomly rotates the map (quarter turns) and mirrors it horizontally and vertically. Combined across our set of training maps, this gave us a large pool of distinct training views from the same underlying maps.
Everything else stayed the same: we kept standard PPO settings, unchanged from the earlier baseline, and changed only the environment. This isolated the effect of augmentation.
Training: one full training run on parallel GPU environments, about 7.5 minutes. Evaluation: five fixed evaluation maps, five episodes each, across three episode lengths (1000 / 3000 / 5000 steps).
3. What we saw — the main numbers
Training converged normally
Final training reward landed at essentially the same level as the earlier baseline. So the model trained fully — it was not undertrained.
Coverage regressed by roughly 1.6 to 3.1 points
| Episode length | Map-aug | Earlier baseline | Prior production | Δ vs baseline |
|---|---|---|---|---|
| @1k | 0.5496 | 0.5656 | 0.5610 | −1.60 pp |
| @3k | 0.8316 | 0.8618 | 0.8618 | −3.02 pp |
| @5k | 0.9104 | 0.9412 | 0.9375 | −3.08 pp |
A warning sign: with augmentation the longest episodes always ran to the full length and never finished early by hitting the coverage threshold, whereas the baseline sometimes finished sooner. The augmented model was simply slower to cover the space.
Behavior shifted — the orientation habit disappeared
The model’s action mix changed sharply. The heavy reliance on a single turn direction — the signature of the earlier reward-shaping behavior — collapsed, dropping severalfold. In its place the model spread across mirrored turns, sideways moves, and far more frequent sensor-scanning pauses.
What happened, in plain terms:
- The model dropped its one-sided turning habit. This confirms that the habit was overfitting to specific map orientations rather than a general strategy.
- It switched to more symmetric, orientation-invariant moves — mirrored turns and sideways motion.
- It scanned far more often, using sensor pauses to compensate during the longer episodes.
Per-map breakdown — symmetric maps held up best
| Map @ 5k | Map-aug | Earlier baseline | Δ |
|---|---|---|---|
| open map | 0.9096 | 0.9294 | −1.98 ✓ smallest drop |
| corridor map | 0.9173 | 0.9365 | −1.92 ✓ smallest drop |
| mixed map A | 0.9200 | 0.9507 | −3.06 |
| mixed map B | 0.9048 | 0.9438 | −3.91 |
| mixed map C | 0.9003 | 0.9456 | −4.54 ✗ worst |
Maps that are already symmetric (open field, corridors) were least affected. Maps with mixed geometry dropped the most. Augmentation removed the orientation-specific patterns that had been working precisely on those mixed maps.
What we concluded
- We expected a clear lift from augmentation; instead we got a regression. The literature did not transfer directly — the classic augmentation work relied on visually distinct training and test data, whereas our maps are semantically similar to each other.
- The combined hope (the earlier reward tweak plus augmentation) was worse even than the prior production model.
- The disappearance of the one-sided turning habit, paired with the coverage drop, told us clearly that the habit had been useful overfitting, not a general mechanism.
- The effect was uneven across maps: symmetric maps were better protected.
The key takeaway
An orientation-invariant model can be worse than a specialized one when the training and test maps come from the same kind of environment. Augmentation pays off mainly when training and test data look genuinely different; ours do not. The generalization gap was already small, so augmentation added more noise than value.
This echoes an earlier finding: a technique borrowed from the literature did not work on our particular kind of task. The pattern keeps holding — methods from papers help only when their underlying assumption matches our setup. Here that assumption (visually distinct training and test data) simply did not apply.
4. Sources / inspiration
- arxiv 1812.02900 — Cobbe et al. 2018 — Quantifying Generalization in RL — the main motivating paper. Lesson: its results are not automatically portable.
- arxiv 2306.16978 — Jonnarth et al. — coverage path planning — the work behind the reward idea we were building on. It confirmed that orientation-specific patterns matter.
- Internal earlier dev-log entries on the reward-shaping baseline and on our first lesson about domain-conditional literature.
5. What’s next
The next experiment shifts toward sim-to-real, by Aleks’s call. A carryover requirement is a dedicated single-sensor failure test, separate from general sensor-noise testing, with noise calibration focused on the rangefinder.
For that work the production model stays the earlier reward-shaping peak, not the augmented one, with the prior production model kept as a secondary baseline.
A few lower-cost ideas remain on the backlog, including a milder version of map augmentation, a short curriculum run, and a longer-horizon hierarchical approach to revisit later.
Glossary
- Data augmentation — generating training-data diversity through transformations such as rotations, mirrors, or noise. For us, it was map rotations and mirrors.
- Generalization gap — the difference between performance on training data and on test data. A large gap suggests overfitting.
- Orientation habit — a strong, one-sided turning preference the model had developed. Augmentation showed it was overfitting to specific map orientations.
- Invariant policy — a model that behaves the same regardless of input orientation. It is often less effective on a specific task than a specialized policy.
- PPO — Proximal Policy Optimization, the reinforcement-learning algorithm used to train the model.