TL;DR
On a representative indoor coverage task in warehouse_v2, the PPO policy trained on 1k episodes beats the classical A* + Frontier baseline by 14.2% on coverage and 22.4% on time-to-95%, at the cost of 3.1% more energy per square meter. On a separate out-of-distribution map (map_04, an open warehouse the policy wasn’t trained on), the policy underperforms the baseline by 8% on coverage — illustrating that RL gains don’t generalize for free.
What we compared
| RL policy (PPO) | Classical baseline (A* + Frontier) | |
|---|---|---|
| Algorithm | PPO, MlpPolicy, γ=0.995, entropy=0.01 | A* planning + Frontier exploration |
| Training | 1k episodes in warehouse_v2, ~12h on RTX 5070 |
n/a (hand-tuned, no training) |
| Observation | 4-stacked TF-Luna + IMU + flow | Same sensor input |
| Action | Normalized velocity (3-DOF) | Same |
| Reward / objective | Coverage bonus + distance penalty + energy penalty | Coverage maximization (greedy) |
Both controllers ran on the same simulated warehouse_v2 map, with identical TF-Luna noise parameters (σ=0.05 m, 100 Hz) and identical iris-claudedrone airframe. Starting position, battery model, and time budget were also held constant.
Results: map_03 (cluttered corridors)
| Metric | RL (PPO) | Classical | Δ |
|---|---|---|---|
| Coverage (final) | 0.948 | 0.806 | +14.2% |
| Time-to-95% coverage | 4 m 12 s | 5 m 24 s | −22.4% |
| Energy per m² | 1.247 | 1.209 | +3.1% |
| Collisions | 0 | 2 | — |
The RL policy wins decisively on coverage and time. The energy penalty is real but small — the policy works the motors slightly harder, which is the price of avoiding the “carousel” pattern (A*'s tendency to revisit already-covered points because the frontier-exploration heuristic doesn’t disincentivize backtracking).
Where the policy wins
Three behaviors emerged during training that the classical baseline doesn’t have:
- Sticky exploration. The policy learns a
rot_-15pattern — slightly biased rotation when entering a new corridor — that handles wall-following better than the baseline’s frontier-frontier-frontier zig-zag. - 2-3 step lookahead. The coverage bonus is smoothed over a 3-step window, which gives PPO incentive to plan ahead. A* + Frontier is greedy by construction.
- Noise tolerance. Under σ=0.20 noise injection (4× the calibrated TF-Luna noise floor), RL coverage drops to 0.91; the classical baseline drops to 0.61. The policy doesn’t fall off a cliff when sensors degrade.
Results: map_04 (open warehouse, out-of-distribution)
| Metric | RL (PPO) | Classical | Δ |
|---|---|---|---|
| Coverage (final) | 0.84 | 0.92 | −8.0% |
| Time-to-95% coverage | DNF | 6 m 18 s | — |
| Energy per m² | 1.103 | 1.084 | +1.7% |
This is where it gets uncomfortable. On map_04 — a large, open warehouse with sparse obstacles — the policy that crushed map_03 underperforms the classical baseline by 8%. The reason is straightforward: map_04 looks nothing like the training distribution. The policy has learned wall-hugging behavior that’s useful in cluttered corridors but actively harmful when the optimal strategy is to fly straight across an open space.
This is the well-known out-of-distribution generalization gap in RL. It’s not a sign that RL doesn’t work; it’s a sign that the training distribution didn’t cover this scenario. The fix is domain randomization during training — varying map geometry, obstacle density, and lighting across episodes — which is the focus of TASK-RL-NEXT-A.
Methodology
- 30 runs per condition (RL / classical × map_03 / map_04), starting from the same five spawn positions, six runs each.
- Coverage measured as the fraction of map cells (10 cm grid) the drone passed within 50 cm of, with 2D LOS to the cell.
- Time-to-95% is the wall-clock time from arm to first reaching 95% coverage; DNF means the run hit the 10-minute budget without reaching it.
- Energy is integrated motor PWM (proxy for battery drain) per m² covered.
- Statistical tests: paired t-test, p < 0.01 for both coverage and time differences on map_03.
The full configuration, dataset, and raw logs are in /rl/dev_log under TASK-RL-EXP-7. Reproducing the result requires the trained checkpoint (also archived under that task) and a working warehouse_v2 setup.
What’s next
- TASK-RL-NEXT-A: noise-aware retraining with σ ∈ {0, 0.1, 0.2, 0.3} interleaved. Hypothesis: implicit regularization closes the out-of-distribution gap on
map_04. - TASK-RL-NEXT-MAP-AUG: training across procedurally-generated maps, not just
warehouse_v2. This is the larger fix for the generalization issue. - Real-drone evaluation: same metrics, same maps, on physical hardware. Pending hardware bring-up (see roadmap).
Where to go next
- RL framework overview — what the policy actually is
- TF-Luna sensor model — what the policy sees
- Last-mile delivery — application where this performance translates to dollars