claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

Crash handling reward

Collision penalty design — terminal vs non-terminal, magnitude calibration.

stabledoc-seoupdated 2026-05-11T00:00:00.000ZClaudeDrone

The collision penalty is a deceptively important reward component. Too small and the policy ignores safety; too large and it becomes risk-averse to the point of never exploring. We settled on −1 per collision, non-terminal episode.

Why non-terminal

A natural first instinct is to terminate the episode on collision (large negative reward → episode ends). The problem with this in our setup:

  • Coverage credit is lost. If the policy crashed at 30% coverage, episodes ending early prevent the credit-assignment chain from working — the policy never sees the long-term consequence of having reached 30% coverage before the crash.
  • PPO advantage estimation degrades. Short episodes give noisy advantage estimates.

So we keep the episode running, reset the drone position to the last safe cell, and apply the −1 penalty without termination.

Magnitude calibration

−1 matches the magnitude of +1 per new cell visited. The implicit message to the policy: a collision costs as much as discovering one new cell. With ~16 cells gained per forward-until-collision call on average, the policy can afford to risk occasional collisions if it discovers enough cells in the process — which matches the desired behavior on the boundary between exploration and caution.

What we tried and rejected

  • −10 per collision — too pessimistic, policy avoided all forward-until-collision actions, coverage tanked.
  • Terminal collision penalty (done=True) — degraded PPO learning (see above).
  • Step-recovery penalty (multiplicative cost on collision) — too complex; didn’t beat the simple version.

Where to go next

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR