Velocity control means the policy emits a normalized three-degree-of-freedom velocity vector — forward speed, lateral speed, and yaw rate — expressed in the drone’s body frame. A downstream publisher node scales this vector into a standard geometry_msgs/Twist message and forwards it to MAVROS as a velocity setpoint. ArduPilot’s state estimator and onboard velocity controller then turn that request into stable motor commands.
This is the default action layer for the platform. It works well for three reasons:
- The hard part of stabilization is already solved. The flight controller’s estimator fuses inertial, rangefinder, and optical-flow data into a reliable velocity estimate, and its velocity controller tracks the requested motion precisely. The policy only needs to decide where to go, not how to keep the drone airborne.
- Sim-to-real transfer is easier. Velocity is a higher-level abstraction than raw motor commands, so the gap between simulated and real-world velocity tracking is much smaller than the gap in low-level motor dynamics.
- Learning is easier to guide. When the action is a velocity, the consequence of a move — getting closer to or further from a goal — is visible almost immediately. With raw motor outputs you would have to wait many control cycles for the airframe to stabilize before the effect of any single action becomes measurable.
Action vector
The continuous action is a three-element vector, each component normalized to the range from -1 to 1: forward velocity, lateral velocity, and yaw rate. The publisher node scales these normalized values into physical units, applying a modest forward and lateral speed envelope and a bounded yaw rate. The conservative speed envelope is deliberate: in cluttered indoor flight, reaction time and predictable motion matter more than peak speed.
Continuous versus discrete action spaces
The same task can be framed with either a continuous or a discrete action space.
A continuous action space lets the policy output a real-valued velocity vector directly, giving fine-grained control over speed and heading. This is the form used for real-drone flight, where smooth, precisely metered motion is important.
A discrete action space instead exposes a small, fixed menu of high-level moves — for example, stepping one grid cell in each cardinal direction, rotating by a fixed angle, holding position to scan, or advancing forward until an obstacle is reached. Under the hood each discrete choice still maps onto the underlying velocity commands, but the policy chooses from a finite set rather than a continuous range. A discrete envelope can simplify exploration and make early training in a grid-based simulator more tractable.
Choosing between them is largely a training-environment decision. The discrete framing is a simulator-side convenience; the real-drone action space remains the continuous velocity vector described above.
Where to go next
- The publisher node that consumes velocity actions and emits the ROS2
Twistmessage - The broader action-space overview for sibling control schemes
- Background on how the flight controller stabilizes the airframe under ROS2 Jazzy
Reinforcement learning on this action space typically uses an on-policy method such as PPO, which pairs naturally with continuous control.