Energy efficiency rewards penalize sustained high-throttle or rapid acceleration, encouraging the policy to move at moderate speeds when possible and to avoid jerky maneuvers. For a real drone with a finite battery, energy efficiency directly maps to flight time — and flight time directly maps to mission usefulness.
Current status: considered but not active in current production reward. The 2D coverage simulator doesn’t model battery state, so an energy penalty wouldn’t have a feedback signal. Becomes relevant when the simulator includes a battery model and when real-drone training begins — see H2 2026 roadmap.
Planned design
energy_step = -k_energy * (commanded_velocity_magnitude)²
Quadratic in velocity matches the real physics — drag scales with v², motor power scales with v²·(thrust requirement). A linear penalty would underweight high-speed motion.
Coefficient k_energy ~ 0.001 — small enough not to override the +1/cell coverage reward, large enough to make the policy notice that 0.5 m/s × 1000 steps is more efficient than 1.0 m/s × 500 steps + lots of stopping.
Connection to the RL vs lawnmower benchmark
The current RL vs lawnmower benchmark reports a 3.1% energy overhead for the RL policy vs the classical baseline. That’s the gap that an explicit energy-efficiency reward would target.
Where to go next
- Reward functions hub — sibling pages
- RL vs lawnmower benchmark — current energy numbers
- Roadmap H2 2026 — when battery-aware training lands