The reward function is where most RL papers spend most of their pages. In practice, our dev-log shows that for indoor coverage, action-space design matters more than reward design — see dev-log 03 (frontier-reward null) and dev-log 07 (multi-cell action step-change). That said, reward design still matters — three pages here cover the main components.
Crash handling — collision penalty, terminal vs non-terminal. Distance penalty — small per-step cost that encourages efficiency. Energy efficiency — motor-throttle penalty for sustained high-power maneuvers.
The baseline reward we settled on:
+1 per new cell visited
+0.1 per scan
-1 per collision (terminal? no — episode continues, but penalty applies)
-0.01 per step
+50 terminal for >95% coverage
This composition is the result of a long sequence of ablations recorded across the RL dev-log. The full justification for each line is split across those entries; the summary above is what currently runs in production.
Contents
Auto-generated from child entries during build (update-indexes.mjs).