The collision penalty is a deceptively important reward component. Too small and the policy ignores safety; too large and it becomes risk-averse to the point of never exploring. We settled on −1 per collision, non-terminal episode.
Why non-terminal
A natural first instinct is to terminate the episode on collision (large negative reward → episode ends). The problem with this in our setup:
- Coverage credit is lost. If the policy crashed at 30% coverage, episodes ending early prevent the credit-assignment chain from working — the policy never sees the long-term consequence of having reached 30% coverage before the crash.
- PPO advantage estimation degrades. Short episodes give noisy advantage estimates.
So we keep the episode running, reset the drone position to the last safe cell, and apply the −1 penalty without termination.
Magnitude calibration
−1 matches the magnitude of +1 per new cell visited. The implicit message to the policy: a collision costs as much as discovering one new cell. With ~16 cells gained per forward-until-collision call on average, the policy can afford to risk occasional collisions if it discovers enough cells in the process — which matches the desired behavior on the boundary between exploration and caution.
What we tried and rejected
- −10 per collision — too pessimistic, policy avoided all forward-until-collision actions, coverage tanked.
- Terminal collision penalty (
done=True) — degraded PPO learning (see above). - Step-recovery penalty (multiplicative cost on collision) — too complex; didn’t beat the simple version.
Where to go next
- Reward functions hub — sibling pages
- Multi-cell forward (dev-log 07) — where the −1 calibration was first validated
- Distance penalty — companion per-step cost