Direct PWM control means the RL policy emits motor PWM signals (~1000-2000 µs per motor) directly, bypassing the flight controller’s stabilization layer. Conceptually attractive because it gives the policy maximum control flexibility; practically problematic because it puts low-level flight stability on the policy’s shoulders, and PPO doesn’t learn fast stabilizing controllers easily.
Current status: considered, not adopted. We use velocity control instead, which keeps ArduPilot’s EKF-fused stabilization in the loop. Direct PWM might become interesting for advanced research (aggressive maneuvers, agility tasks) but isn’t appropriate for the indoor-coverage task that drives current development.
Why we didn’t pick this
- Stability has to be learned from scratch. ArduPilot’s EKF + control loop solves this; a PPO policy emitting raw PWM at ~10 Hz can’t keep up with the ~400 Hz rate flight controllers actually run at.
- Sample efficiency for stability learning is poor. The drone would crash thousands of times before learning to hover.
- Sim-to-real transfer is harder. Motor dynamics differ subtly between sim and real (see custom motor physics). With PWM, the policy depends on those dynamics being right; with velocity, the FC absorbs the mismatch.
When this becomes interesting
- Agility / racing tasks where precise low-level control beats high-level setpoint commands.
- Hybrid hierarchies where a high-level RL planner picks waypoints and a low-level RL controller emits PWM to track them.
Where to go next
- Velocity control — what we use
- Action space hub — sibling pages
action_publishernode — the consumer of the policy’s output