claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

20 — Map augmentation removed an orientation trick, and overall coverage dropped

Rotating and mirroring training maps removed an orientation-specific habit, and indoor coverage regressed by a few points.

stablerl-labupdated 2026-05-11T00:00:00.000ZClaudeDroneRLDevLog

What this is: we added map rotations and mirrors to training — each map could appear in any of its rotated and mirrored orientations. The idea: the model should learn to work on any orientation of a map rather than a memorized shape. Why: in an earlier experiment a reward-shaping variant gave a small coverage gain. We wanted to see whether map augmentation could amplify that result, since augmentation is a standard generalization trick in the literature.

Date: 10 May 2026 Machine: D2


1. What we wanted to test

Context: in the previous experiment our reward-shaping variant produced a small coverage gain over the earlier production baseline. Looking at how the model behaved, we noticed something odd: it leaned very heavily on one particular turn direction. That could be either:

  • a genuinely smart traversal strategy that the reward had taught, or
  • overfitting to the specific orientations of the training maps.

Hypothesis: if we rotate and mirror the maps during training, the model can no longer lean on a specific orientation. So:

  • If the helpful behavior is real and transferable, coverage should hold steady or improve.
  • If it was overfitting, coverage should fall — but we would at least know we had an honest, orientation-invariant model.

We went in expecting augmentation to help, based on prior generalization work, and treated the result as a genuine test of whether that habit was a real mechanism or just memorization.

2. What we did

We built a new training environment that, on every reset, randomly rotates the map (quarter turns) and mirrors it horizontally and vertically. Combined across our set of training maps, this gave us a large pool of distinct training views from the same underlying maps.

Everything else stayed the same: we kept standard PPO settings, unchanged from the earlier baseline, and changed only the environment. This isolated the effect of augmentation.

Training: one full training run on parallel GPU environments, about 7.5 minutes. Evaluation: five fixed evaluation maps, five episodes each, across three episode lengths (1000 / 3000 / 5000 steps).

3. What we saw — the main numbers

Training converged normally

Final training reward landed at essentially the same level as the earlier baseline. So the model trained fully — it was not undertrained.

Coverage regressed by roughly 1.6 to 3.1 points

Episode length Map-aug Earlier baseline Prior production Δ vs baseline
@1k 0.5496 0.5656 0.5610 −1.60 pp
@3k 0.8316 0.8618 0.8618 −3.02 pp
@5k 0.9104 0.9412 0.9375 −3.08 pp

A warning sign: with augmentation the longest episodes always ran to the full length and never finished early by hitting the coverage threshold, whereas the baseline sometimes finished sooner. The augmented model was simply slower to cover the space.

Behavior shifted — the orientation habit disappeared

The model’s action mix changed sharply. The heavy reliance on a single turn direction — the signature of the earlier reward-shaping behavior — collapsed, dropping severalfold. In its place the model spread across mirrored turns, sideways moves, and far more frequent sensor-scanning pauses.

What happened, in plain terms:

  1. The model dropped its one-sided turning habit. This confirms that the habit was overfitting to specific map orientations rather than a general strategy.
  2. It switched to more symmetric, orientation-invariant moves — mirrored turns and sideways motion.
  3. It scanned far more often, using sensor pauses to compensate during the longer episodes.

Per-map breakdown — symmetric maps held up best

Map @ 5k Map-aug Earlier baseline Δ
open map 0.9096 0.9294 −1.98 ✓ smallest drop
corridor map 0.9173 0.9365 −1.92 ✓ smallest drop
mixed map A 0.9200 0.9507 −3.06
mixed map B 0.9048 0.9438 −3.91
mixed map C 0.9003 0.9456 −4.54 ✗ worst

Maps that are already symmetric (open field, corridors) were least affected. Maps with mixed geometry dropped the most. Augmentation removed the orientation-specific patterns that had been working precisely on those mixed maps.

What we concluded

  • We expected a clear lift from augmentation; instead we got a regression. The literature did not transfer directly — the classic augmentation work relied on visually distinct training and test data, whereas our maps are semantically similar to each other.
  • The combined hope (the earlier reward tweak plus augmentation) was worse even than the prior production model.
  • The disappearance of the one-sided turning habit, paired with the coverage drop, told us clearly that the habit had been useful overfitting, not a general mechanism.
  • The effect was uneven across maps: symmetric maps were better protected.

The key takeaway

An orientation-invariant model can be worse than a specialized one when the training and test maps come from the same kind of environment. Augmentation pays off mainly when training and test data look genuinely different; ours do not. The generalization gap was already small, so augmentation added more noise than value.

This echoes an earlier finding: a technique borrowed from the literature did not work on our particular kind of task. The pattern keeps holding — methods from papers help only when their underlying assumption matches our setup. Here that assumption (visually distinct training and test data) simply did not apply.

4. Sources / inspiration

5. What’s next

The next experiment shifts toward sim-to-real, by Aleks’s call. A carryover requirement is a dedicated single-sensor failure test, separate from general sensor-noise testing, with noise calibration focused on the rangefinder.

For that work the production model stays the earlier reward-shaping peak, not the augmented one, with the prior production model kept as a secondary baseline.

A few lower-cost ideas remain on the backlog, including a milder version of map augmentation, a short curriculum run, and a longer-horizon hierarchical approach to revisit later.

Glossary

  • Data augmentation — generating training-data diversity through transformations such as rotations, mirrors, or noise. For us, it was map rotations and mirrors.
  • Generalization gap — the difference between performance on training data and on test data. A large gap suggests overfitting.
  • Orientation habit — a strong, one-sided turning preference the model had developed. Augmentation showed it was overfitting to specific map orientations.
  • Invariant policy — a model that behaves the same regardless of input orientation. It is often less effective on a specific task than a specialized policy.
  • PPO — Proximal Policy Optimization, the reinforcement-learning algorithm used to train the model.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR