The per-step distance penalty is the smallest reward component (−0.01 per step) but does important work: it gives the policy a reason to finish episodes quickly rather than wandering aimlessly. Without it, a policy that scans-and-rotates forever would have similar reward to one that actually covers the map.
Calibration logic
Per-step penalty: −0.01.
Per-cell gain: +1.
So the policy needs to cover at least one new cell per 100 steps to break even. In practice it covers far more than that during productive phases (especially during forward-until-collision action 7), so the penalty is small enough not to discourage exploration but large enough to prevent infinite scanning.
What we tried and rejected
- −0.001 per step (10× weaker): policy lingered in already-visited areas waiting for new state. Coverage dropped.
- −0.1 per step (10× stronger): policy became frantic, took risky shortcuts, collided more.
Note on time-penalty vs distance-penalty
We call this a “distance” penalty by convention, but it’s actually a time penalty — applied per step regardless of motion. A pure distance penalty (per cell moved) was considered and rejected: it disincentivizes the long-traversal action 7 too aggressively. The current per-step formulation lets action 7 cover many cells “for free” on a single step, which preserves the action’s effectiveness.
Where to go next
- Reward functions hub — sibling pages
- Crash handling — companion safety penalty
- Energy efficiency — adjacent efficiency signal