Train a flight policy in simulation and it learns the world through perfect eyes. Every distance reading is exact, every sensor always answers, and walls sit precisely where the map says they are. Then you bolt a five-dollar time-of-flight (ToF) sensor onto a real drone, and that tidy world evaporates. The gap between “what the policy practiced on” and “what the hardware actually reports” is the sim-to-real gap, and on cheap distance sensors it is where a lot of otherwise-good policies quietly fall apart.
This article is about what survives that jump and what bites you — drawn from a noise sweep our reinforcement-learning team ran on real sensor models, and from a sensor-modeling study of the TF-Luna and VL53L0X families.
Why a clean policy breaks on a dirty sensor
A simulator is a comfortable place to learn. Distances come back to the millimetre, instantly, every single step. A policy trained there can grow dependent on that precision — it learns to trust the number in front of it completely, because in sim the number is never wrong.
Real ToF sensors do not extend that courtesy. A VL53L0X carries roughly ±3% error. A TF-Luna is accurate to about ±6 cm at close-to-mid range but can drift by several centimetres further out. Both drop readings entirely under bright light, reflective surfaces, or low-contrast targets, and both can hiccup when mounted on a moving servo. The honest version of the sensor is noisy, range-limited, and occasionally silent.
A policy that never met that messiness treats each ugly reading as a contradiction of everything it knows. The result is brittle, twitchy behavior on hardware — exactly the failure the lab set out to measure.
Domain randomization: practice in the mess
The fix is delightfully simple in concept: stop training in a pristine world. Domain randomization means deliberately roughing up the simulation — injecting realistic noise, offsets, and dropouts — so the policy learns the task rather than memorizing one clean sensor.
Think of a musician who only ever rehearses in a silent, acoustically perfect room, then has to perform at a noisy outdoor festival. The notes are the same; the conditions are a shock. A musician who practiced with crowd noise, wind, and a slightly out-of-tune monitor walks on stage unbothered. Domain randomization is making the drone rehearse with the festival already running.
In practice that means adding Gaussian noise to simulated distance readings, nudging the sensor’s mounting angle and calibration offset a little each episode, and occasionally throwing away a reading to mimic a dropout. The policy stops assuming the number is gospel and learns to fly well even when it is a few centimetres off.
How much noise actually matters
The encouraging news from the sweep: for realistic noise, the sim-to-real transition is barely a problem at all. Real VL53L0X and TF-Luna sensors operate around σ ≈ 0.03 (about 3% of range). In testing, policies lost essentially nothing up to σ ≈ 0.20 — roughly seven times worse than reality — and only started to degrade meaningfully past σ ≈ 0.5, which is “comically bad sensor” territory.
For training, a sweet spot emerges around σ ≈ 0.05 m (5 cm). That sits right on top of real-world TF-Luna variance: clean enough that the policy can still learn, noisy enough that it never overfits to a perfect sensor. Going much higher (σ ≈ 0.20) over-randomizes — the world becomes so noisy that learning itself suffers. The shape to remember: a little noise is protective, too much is corrosive.
The trait that really bites: a missing forward reading
Noise is the gentle part of the story. The sharp edge is dropouts and dead channels. A policy calibrated on a full ring of sensors reads a single stuck-at-zero channel as an impossible, out-of-distribution signal — and that can hurt more than losing several sensors at once.
The lab found the forward-facing VL53L0X to be the one mission-critical eye: the policy’s long, confident moves depend on knowing the distance straight ahead. Lose that one reading and the drone can no longer tell whether a wall is one cell away or ten, so it stops committing to runs, crawls cell by cell, and rescans constantly. The lesson is blunt: model dropouts in sim, and on hardware, protect the forward axis — duplicate it, or detect a stuck reading and fail over to a backup.
A sim-to-real checklist for ToF
- Inject Gaussian distance noise in sim. Start near σ ≈ 0.05 m to match real TF-Luna / VL53L0X variance. Don’t push past ~0.2 — you’ll starve learning.
- Randomize mounting and calibration. Vary sensor-frame rotation (~±2°) and a Z offset (~±10 cm) per episode so the policy tolerates a real-world mount that’s never perfect.
- Simulate dropouts. Drop a small fraction of readings per step to mimic servo motion, bus glitches, and overexposure — this is what makes a policy robust, not just noise.
- Model per-sensor failure, especially forward. Train with the occasional dead channel so a single zeroed reading isn’t an alien signal at deploy time.
- Match real range limits. A TF-Luna reads to ~8 m, a VL53L0X far less — clamp sim ranges so the policy never relies on distances the hardware can’t deliver.
- Trust the noise margin. Realistic ToF noise sits ~10× below where policies degrade; don’t over-engineer for noise you’ll never see. Spend the effort on dropouts and the forward sensor instead.
The takeaway
Cheap ToF sensors transfer better than their price suggests — if you train for the world they actually live in. Distance noise at realistic levels is a non-issue with a modest amount of domain randomization. The real hazard is the silent channel: a dropout or dead forward sensor that a clean-trained policy has no idea how to handle. Rehearse in the mess, protect the forward axis, and the jump from Gazebo to hardware stops being a leap of faith.
Cross-references: dev-log 21 — sim-to-real noise sweep and the ToF sim-to-real reference.