claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

Distance penalty

Per-step cost that encourages efficient traversal.

stabledoc-seoupdated 2026-05-11T00:00:00.000ZClaudeDrone

The per-step distance penalty is the smallest reward component (−0.01 per step) but does important work: it gives the policy a reason to finish episodes quickly rather than wandering aimlessly. Without it, a policy that scans-and-rotates forever would have similar reward to one that actually covers the map.

Calibration logic

Per-step penalty: −0.01.

Per-cell gain: +1.

So the policy needs to cover at least one new cell per 100 steps to break even. In practice it covers far more than that during productive phases (especially during forward-until-collision action 7), so the penalty is small enough not to discourage exploration but large enough to prevent infinite scanning.

What we tried and rejected

  • −0.001 per step (10× weaker): policy lingered in already-visited areas waiting for new state. Coverage dropped.
  • −0.1 per step (10× stronger): policy became frantic, took risky shortcuts, collided more.

Note on time-penalty vs distance-penalty

We call this a “distance” penalty by convention, but it’s actually a time penalty — applied per step regardless of motion. A pure distance penalty (per cell moved) was considered and rejected: it disincentivizes the long-traversal action 7 too aggressively. The current per-step formulation lets action 7 cover many cells “for free” on a single step, which preserves the action’s effectiveness.

Where to go next

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR