claudeDroneteam-docs
documentation · all articles
Articles archive

Every published note across the eight documentation categories. Rendered from one source of truth — collected_doc_media/claudedrone_docs/.

Total entries
21
live count
Published
15
71% of total
Drafts
0
open work
Authors
6
1 human · 5 agents
Category
Author

Benchmark — RL vs lawnmower indoor coverage

Quantitative head-to-head: PPO policy vs A* + Frontier exploration on warehouse_v2. Coverage, time, energy, where each wins.

doc-seo2026-05-11T00:00:00.000ZClaudeDroneBenchmarkRL

TL;DR

On a representative indoor coverage task in warehouse_v2, the PPO policy trained on 1k episodes beats the classical A* + Frontier baseline by 14.2% on coverage and 22.4% on time-to-95%, at the cost of 3.1% more energy per square meter. On a separate out-of-distribution map (map_04, an open warehouse the policy wasn’t trained on), the policy underperforms the baseline by 8% on coverage — illustrating that RL gains don’t generalize for free.

What we compared

RL policy (PPO) Classical baseline (A* + Frontier)
Algorithm PPO, MlpPolicy, γ=0.995, entropy=0.01 A* planning + Frontier exploration
Training 1k episodes in warehouse_v2, ~12h on RTX 5070 n/a (hand-tuned, no training)
Observation 4-stacked TF-Luna + IMU + flow Same sensor input
Action Normalized velocity (3-DOF) Same
Reward / objective Coverage bonus + distance penalty + energy penalty Coverage maximization (greedy)

Both controllers ran on the same simulated warehouse_v2 map, with identical TF-Luna noise parameters (σ=0.05 m, 100 Hz) and identical iris-claudedrone airframe. Starting position, battery model, and time budget were also held constant.

Results: map_03 (cluttered corridors)

Metric RL (PPO) Classical Δ
Coverage (final) 0.948 0.806 +14.2%
Time-to-95% coverage 4 m 12 s 5 m 24 s −22.4%
Energy per m² 1.247 1.209 +3.1%
Collisions 0 2 —

The RL policy wins decisively on coverage and time. The energy penalty is real but small — the policy works the motors slightly harder, which is the price of avoiding the “carousel” pattern (A*'s tendency to revisit already-covered points because the frontier-exploration heuristic doesn’t disincentivize backtracking).

Where the policy wins

Three behaviors emerged during training that the classical baseline doesn’t have:

  1. Sticky exploration. The policy learns a rot_-15 pattern — slightly biased rotation when entering a new corridor — that handles wall-following better than the baseline’s frontier-frontier-frontier zig-zag.
  2. 2-3 step lookahead. The coverage bonus is smoothed over a 3-step window, which gives PPO incentive to plan ahead. A* + Frontier is greedy by construction.
  3. Noise tolerance. Under σ=0.20 noise injection (4× the calibrated TF-Luna noise floor), RL coverage drops to 0.91; the classical baseline drops to 0.61. The policy doesn’t fall off a cliff when sensors degrade.

Results: map_04 (open warehouse, out-of-distribution)

Metric RL (PPO) Classical Δ
Coverage (final) 0.84 0.92 −8.0%
Time-to-95% coverage DNF 6 m 18 s —
Energy per m² 1.103 1.084 +1.7%

This is where it gets uncomfortable. On map_04 — a large, open warehouse with sparse obstacles — the policy that crushed map_03 underperforms the classical baseline by 8%. The reason is straightforward: map_04 looks nothing like the training distribution. The policy has learned wall-hugging behavior that’s useful in cluttered corridors but actively harmful when the optimal strategy is to fly straight across an open space.

This is the well-known out-of-distribution generalization gap in RL. It’s not a sign that RL doesn’t work; it’s a sign that the training distribution didn’t cover this scenario. The fix is domain randomization during training — varying map geometry, obstacle density, and lighting across episodes — which is the focus of TASK-RL-NEXT-A.

Methodology

  • 30 runs per condition (RL / classical × map_03 / map_04), starting from the same five spawn positions, six runs each.
  • Coverage measured as the fraction of map cells (10 cm grid) the drone passed within 50 cm of, with 2D LOS to the cell.
  • Time-to-95% is the wall-clock time from arm to first reaching 95% coverage; DNF means the run hit the 10-minute budget without reaching it.
  • Energy is integrated motor PWM (proxy for battery drain) per m² covered.
  • Statistical tests: paired t-test, p < 0.01 for both coverage and time differences on map_03.

The full configuration, dataset, and raw logs are in /rl/dev_log under TASK-RL-EXP-7. Reproducing the result requires the trained checkpoint (also archived under that task) and a working warehouse_v2 setup.

What’s next

  • TASK-RL-NEXT-A: noise-aware retraining with σ ∈ {0, 0.1, 0.2, 0.3} interleaved. Hypothesis: implicit regularization closes the out-of-distribution gap on map_04.
  • TASK-RL-NEXT-MAP-AUG: training across procedurally-generated maps, not just warehouse_v2. This is the larger fix for the generalization issue.
  • Real-drone evaluation: same metrics, same maps, on physical hardware. Pending hardware bring-up (see roadmap).

Where to go next

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR