Soft Actor-Critic (SAC) is the natural alternative to PPO when the action space is continuous. It’s typically more sample-efficient than PPO on robotic-control tasks and uses entropy regularization explicitly in the objective.
Current status: not in production. Our current 2D coverage simulator uses a discrete action space (eight moves), where PPO’s strengths fit better than SAC’s. SAC would become relevant if/when we switch to continuous velocity control for the inner-loop policy in a hierarchical setup, or for low-level real-drone control.
Why we haven’t pursued SAC
- Discrete action space dominates current work. SAC is fundamentally a continuous-action algorithm; the discrete variants (SAC-Discrete) exist but lose the advantages.
- PPO is well-tuned and stable. Our PPO baseline produces predictable results across the 24 dev-log experiments. Switching algorithms mid-stream would add a confound.
- The breakthroughs came from representation, not algorithm. The biggest wins (multi-cell forward action, observation engineering) would transfer to SAC too — and PPO already exploited them.
When SAC would become interesting
- Continuous-velocity inner-loop policy in a hierarchical RL setup, where a top-level PPO picks regions and a bottom-level SAC controls smooth velocity tracking.
- Real-drone fine-tuning where sample efficiency matters more than sim-only throughput.
Where to go next
- PPO hyperparameters — what we use today
- Algorithms hub — sibling pages
- Roadmap H2 2026 — when hierarchical work might land