claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

Energy efficiency reward

Penalty for sustained high-power maneuvers — encourages battery-aware policies.

stabledoc-seoupdated 2026-05-11T00:00:00.000ZClaudeDrone

Energy efficiency rewards penalize sustained high-throttle or rapid acceleration, encouraging the policy to move at moderate speeds when possible and to avoid jerky maneuvers. For a real drone with a finite battery, energy efficiency directly maps to flight time — and flight time directly maps to mission usefulness.

Current status: considered but not active in current production reward. The 2D coverage simulator doesn’t model battery state, so an energy penalty wouldn’t have a feedback signal. Becomes relevant when the simulator includes a battery model and when real-drone training begins — see H2 2026 roadmap.

Planned design

energy_step = -k_energy * (commanded_velocity_magnitude)²

Quadratic in velocity matches the real physics — drag scales with v², motor power scales with v²·(thrust requirement). A linear penalty would underweight high-speed motion.

Coefficient k_energy ~ 0.001 — small enough not to override the +1/cell coverage reward, large enough to make the policy notice that 0.5 m/s × 1000 steps is more efficient than 1.0 m/s × 500 steps + lots of stopping.

Connection to the RL vs lawnmower benchmark

The current RL vs lawnmower benchmark reports a 3.1% energy overhead for the RL policy vs the classical baseline. That’s the gap that an explicit energy-efficiency reward would target.

Where to go next

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR