Two pages here. PPO hyperparameters documents the specific PPO configuration we use — gamma, learning rate, entropy coefficient, rollout size — and the rationale behind each. SAC experiments covers Soft Actor-Critic as an alternative we’ve considered but not pursued in production.
PPO is the workhorse. It’s stable, well-implemented in Stable-Baselines3, and works well with discrete action spaces (which is what our current 2D coverage env uses). The full experiment chain — bootstrap through 24 dev-log entries — all uses PPO with minor variations. The big wins came from changes to action space (multi-cell forward) and observation engineering, not from algorithm swaps. That’s a useful pattern to internalize: for indoor coverage, the algorithm matters less than the representation.
Contents
Auto-generated from child entries during build (update-indexes.mjs).