The action space defines what the policy can do. Picking the right action space is one of the highest-leverage choices in an RL system — the multi-cell forward action discovery in EXP-7 showed that the right action representation can be worth more than dozens of reward-shaping experiments combined.
Two pages here cover the alternatives. Velocity control is the action layer we currently use — the policy emits a normalized 3-DOF velocity command (forward/lateral/yaw), which the action_publisher node translates into a MAVROS cmd_vel. Direct PWM control is an alternative we’ve considered but not pursued — sending motor PWM signals directly bypasses the flight controller, which gives maximum flexibility but loses ArduPilot’s IMU-fused stabilization. For an indoor drone where stable hovering matters, velocity control is the better choice.
The full set of actions used in training also includes the multi-cell forward macro from EXP-7 — a single action that translates into many environment steps. That’s documented in the dev-log entry that introduced it.
Contents
Auto-generated from child entries during build (update-indexes.mjs).