claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

Reward functions

What we pay the policy for — coverage, distance penalties, crash penalties, energy efficiency.

draftdoc-seoupdated 2026-05-11T00:00:00.000ZClaudeDrone

The reward function is where most RL papers spend most of their pages. In practice, our dev-log shows that for indoor coverage, action-space design matters more than reward design — see dev-log 03 (frontier-reward null) and dev-log 07 (multi-cell action step-change). That said, reward design still matters — three pages here cover the main components.

Crash handling — collision penalty, terminal vs non-terminal. Distance penalty — small per-step cost that encourages efficiency. Energy efficiency — motor-throttle penalty for sustained high-power maneuvers.

The baseline reward we settled on:

+1 per new cell visited
+0.1 per scan
-1 per collision (terminal? no — episode continues, but penalty applies)
-0.01 per step
+50 terminal for >95% coverage

This composition is the result of a long sequence of ablations recorded across the RL dev-log. The full justification for each line is split across those entries; the summary above is what currently runs in production.

Contents

Auto-generated from child entries during build (update-indexes.mjs).

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR