What this is: We tried sim-to-real domain randomization — adding sensor noise, delays, and servo inaccuracy during training. The result surprised us: the model trained without domain randomization was already naturally robust to our noise, while the model trained with it got noticeably worse in clean conditions. Why it’s here: To explain why we are not continuing simulation-side experiments and are moving to real hardware instead. Our simulator has a structural limitation that makes a sim-to-real noise test uninformative.
Date: 2026-05-19 Ticket: TASK-RL-DR (sim-to-real domain randomization — the last simulation step before real hardware)
Glossary
- Domain randomization (DR): A training trick where you sprinkle random noise into the drone’s sensors and actions during learning, so the model gets used to imperfect conditions. The hope is a policy that survives the jump to real hardware. Think of it like practicing a free throw with a slightly deflated ball — if you can sink it under messy conditions, the real game feels easy.
- Clean environment: A simulation with perfect sensors and no noise. This is our standard setup for measuring how good a model is.
- Noisy environment: A simulation with noise added, meant to imitate the imperfections of real hardware.
- VL53L0X: A time-of-flight distance sensor. The drone carries six of them, one facing each side.
- TF-Luna: A longer-range time-of-flight sensor mounted on a servo so it can swivel and scan.
- Servo backlash: The mechanical slack in a servo. Command it to “turn 15°” and the real turn lands somewhere around 12–18° because of the play in the gears. Like the dead zone in an old steering wheel before the car actually reacts.
- Generalization gap: The drop in performance between clean and noisy conditions — basically, how much the model falls apart once the world stops being perfect.
- Coverage (cov@N): The fraction of the room the drone has explored after N steps. cov@5k means coverage after 5,000 steps.
1. What we expected
A real drone will not have perfect senses. We assumed three sources of imperfection:
- The VL53L0X sensors report distance with about ±10% error (roughly Gaussian).
- The TF-Luna has a read delay of 50–100 ms — about one control step’s worth of lag.
- The servo has roughly ±5° of backlash.
Our working idea was simple: if we train the model with this noise baked in, it should be tougher on real hardware. The expected price was a small dip in coverage under ideal conditions — a fair trade, we thought.
2. What we did
- Built
Drone2DEnvDR, a version of the environment that injects sensor noise, read delay, and servo backlash. It is parameterized, so the noise can be switched off entirely. - Trained a model on this noisy version for one million steps, reusing the same setup as SWEEP-02 (our best model to date). Training took about 6.6 minutes.
- Ran a 3×3 evaluation matrix to see how each model behaved in each world:
- SWEEP-02 (no DR) × clean — our baseline.
- SWEEP-02 × noisy — how much it degrades under noise.
- DR × clean — how much it gives up in ideal conditions (the cost of training with DR).
- DR × noisy — our intended deployment target.
3. What we got — two surprises
The numbers
| Condition | cov@1k | cov@3k | cov@5k |
|---|---|---|---|
| SWEEP-02 / clean | 61.9% | 91.8% | 95.0% |
| SWEEP-02 / noisy | 62.8% | 92.6% | 95.0% |
| DR / clean | 28.4% ⚠ | 59.1% ⚠ | 80.1% ⚠ |
| DR / noisy | 62.5% | 91.3% | 95.0% |
Surprise 1: SWEEP-02, with no DR at all, is already robust
The generalization gap for SWEEP-02 — the change from clean to noisy — was one percentage point or less at every horizon. On some runs it was even slightly better under noise. In plain terms: our sensor-noise model has essentially no effect on how the policy performs.
Surprise 2: DR badly hurt the clean case
The DR-trained model, run on the clean environment, dropped by 33 percentage points at 1,000 steps (28.4% versus 61.9%). That is an enormous regression. Under noise it was fine (62.5%), but no better than SWEEP-02 had been.
So the verdict was the opposite of what we expected: DR did not help, and it actively hurt.
4. Why this happened — the key insight
Digging through the environment code explained everything.
- Collision detection in the environment reads the map directly from the ground-truth grid (
grid[y, x]). It does not consult the sensors at all. - The visited grid (64×64) is the main thing the model looks at (a CNN runs over it). This grid is the environment’s internal memory — it updates only when the drone actually visits a cell, and it never depends on noisy sensor readings.
- The “forward until collision” action marches cell by cell along the grid. It does not use sensors to decide when to stop.
So the sensor noise only touches what the policy can see (the distance values in its observation). But the policy mostly steers by the visited grid — and that grid is noise-free. As a result, the policy simply ignores the sensor noise.
Training with DR pushed the policy to be cautious in noisy conditions: it learned to trust its heading less (because of servo backlash) and to trust the forward-until-collision move less (because of distance noise). In clean conditions all of that caution is dead weight, and the 33-point drop is the bill for it.
The short version: we added noise to a part of the system that was never the limiting factor, and the model paid for caution it didn’t need.
5. What this means for production readiness
How Aleks framed it: this was meant to be “the last step before we can honestly say the model is ready for real hardware.”
The honest answer: our current simulator is not a valid sim-to-real test for sensor noise. The environment’s physics do not depend on the sensors at all — collisions are decided by a ground-truth grid lookup. A model that learned to play well in this simulation is robust to sensor noise by construction, but that robustness is an artifact of the simulator, not a real property of the policy.
To actually judge readiness, we need one of two paths:
Option A — rewrite the environment so that collisions are decided from sensor readings, where false positives and false negatives become possible. That is roughly 4–6 hours of work. It would be a much better test — but it is still a simulation.
Option B — real hardware (the path Aleks chose). Port the model onto an ESP32 with real sensors on a real arena. This is the only true source of truth.
So the DR line of work is closed with a clear lesson rather than a fix. SWEEP-02 stands as the final pre-hardware model, and the next session moves to real hardware deployment.
Related files
envs/drone_2d_env_dr.py— the DR environment (new)experiments/TASK-RL-DR/summary.md— the detailed report