What this is: we checked how well our best model and our previous one survive realistic sensor noise. We ran 30 experiments with different noise types, without any retraining. Why: before deploying to a real drone we need to know how much noise a model can tolerate, and which sensor failure would be catastrophic.
Date: 10 May 2026 Machine: D2
1. What we wanted to check
Context: our models learn in an idealized simulation with exact distances and exact position. Real hardware never behaves that cleanly. A VL53L0X has roughly ±3% error, a TF-Luna can be off by several centimetres at long range, the IMU drifts, and a sensor can fail outright mid-flight. We wanted to understand how each of these conditions affects a trained policy.
We went in expecting that moderate noise might degrade the model quickly, that retraining with noise would likely help, and — based on an earlier sensor-failure study — that losing a single sensor might be worse than losing many. We also wanted to compare the robustness of our two best models head to head.
2. What we did
We built a noisy simulation environment that layers three independent kinds of disturbance on top of the normal world:
- Gaussian noise on the distance readings, swept across a wide range to imitate real sensor noise.
- Random errors in the visited-area map, applied with increasing probability per step to imitate localization (SLAM) drift.
- Total failure of a single sensor, applied to each sensor in turn to imitate a hardware break.
We then ran 30 experiments across our two best models, five maps, and several episodes each, evaluating without any retraining.
3. What we saw — the main numbers
ToF sensor noise — the threshold where the model starts to break
| σ (noise) | Best model | Previous model | Δ best | Δ previous |
|---|---|---|---|---|
| 0.0 (no noise) | 0.9384 | 0.9377 | — | — |
| 0.01 (3% — realistic VL53) | 0.9381 | 0.9414 | within noise | within noise |
| 0.05 (5% — typical cheap ToF) | 0.9394 | 0.9368 | within noise | within noise |
| 0.10 (10% — noisy sensors) | 0.9354 | 0.9274 | −0.30 pp | −1.03 pp |
| 0.20 (20% — poor sensors) | 0.9281 | 0.9216 | −1.03 pp | −1.61 pp |
| 0.40 (unrealistically bad) | 0.8949 | 0.8697 | −4.35 pp | −6.79 pp |
| 0.60 | 0.8420 | 0.8037 | −9.64 pp | −13.39 pp |
| 0.80 | 0.7816 | 0.7353 | −15.68 pp | −20.24 pp |
The headline: up to σ=0.20 (20% noise) both models lose almost nothing. Real-world VL53L0X and TF-Luna sensors operate around σ≈0.03 — roughly seven times below where degradation begins. Along this axis, the sim-to-real transition is simply not a problem.
The point where the drop exceeds 10 pp lands near σ≈0.55 for the best model and σ≈0.50 for the previous one. The best model stays about 2–3 pp ahead at every noise level — a welcome bit of extra robustness.
Visited map — almost completely robust
| p (per-step error probability) | Best model | Previous model |
|---|---|---|
| 0.01 (1% — typical SLAM) | parity | parity |
| 0.05 | parity | parity |
| 0.20 (20% per step) | parity | parity |
| 0.50 (50% — every other step the map drifts) | parity | parity |
Every value stays within ±0.5 pp. The model is fully robust to localization errors, so cheap SLAM or drift-prone odometry is fine in practice.
Single-sensor failure — the real hazard
| Broken sensor | Δ coverage | Severity |
|---|---|---|
| vl53_0 (forward) | −79.31 pp | catastrophic |
| vl53_1 (60° right) | −5.24 | moderate |
| vl53_2 (120°, back-right) | −5.28 | moderate |
| vl53_3 (180°, straight back) | −2.48 | within noise |
| vl53_4 (240°, back-left) | −2.32 | within noise |
| vl53_5 (300°, 60° left) | −5.15 | moderate |
| tf_luna (long range, on servo) | −1.46 | within noise |
The forward VL53L0X turns out to be the only mission-critical sensor. Without it the agent essentially collapses, because the model’s main long-reach forward move depends on knowing the distance straight ahead. With no forward reading the agent can’t tell whether a wall is one cell away or ten, so it stops making long runs, crawls one cell at a time, and constantly rescans.
This echoes the earlier “distortion is worse than absence” observation. Zeroing only the forward sensor cost about −79 pp, whereas zeroing all six body-mounted VL53L0X sensors (with the TF-Luna still working) cost about −58.8 pp, and zeroing all seven ToF sensors cost about −55.4 pp. A single broken sensor is worse than losing all of them, because the model is calibrated on the full set of distances and a single zeroed channel looks like an out-of-distribution signal.
Main takeaways
- Strong degradation under moderate noise only appears once σ climbs above about 0.55 — but real-world σ is far smaller (≈0.03), so on hardware everything holds up.
- The model is robust out of the box for realistic noise: the sim-to-real gap is not the main concern here.
- Losing a single sensor really is worse than losing all of them — the earlier pattern clearly transfers to live failures.
- Our best model is consistently 2–3 pp more robust than the previous one, which is one more reason to keep it as the production choice.
- The retrain-with-noise question was deferred to a later session.
4. Hardware recommendations (for the real drone)
- The forward sensor matters most. Either duplicate it with two forward-facing sensors, or detect a “stuck-at-zero” reading and switch to a backup.
- Sensor noise is not scary. Noise up to σ≈0.4 is tolerable, and real ToF sensors at σ≈0.03 sit roughly 10× below that.
- Localization can drift. SLAM drift up to 50% per step is fine.
5. What’s next
When we continue, we plan to explore retraining the policy with noise injected during training to see whether it raises the tolerance threshold, and retraining with occasional simulated sensor failures to soften the forward-sensor catastrophe. There is also a queue of follow-up training experiments aimed at closing the remaining gap to the lawnmower baseline.
Glossary
- Sim-to-real gap — the difference between performance in simulation and on real hardware.
- σ (sigma) — the standard deviation of the noise; σ=0.05 means 5% of the range.
- vl53_0…vl53_5 — our six VL53L0X ToF sensors mounted on the body at 60° intervals.
- tf_luna — a long-range ToF sensor mounted on a rotating servo.
- Distortion > absence — the pattern that one broken sensor can hurt more than all sensors being lost at once.