claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

TensorBoard analysis

Reading PPO training curves — what healthy looks like, what regressions look like.

stabledoc-seoupdated 2026-05-11T00:00:00.000ZClaudeDrone

TensorBoard is the main diagnostic tool during training. It shows you in near-real-time whether the policy is learning, whether the run is degenerate, and whether something is broken. This page covers what to watch for.

Healthy training curves

For a well-tuned PPO run on our coverage task, you should see:

  • ep_rew_mean climbing from ~160 (random) to ~2000 (converged) over 1M steps. Some noise is normal; sudden drops are a red flag.
  • ep_len_mean stable at 1000 (max episode length) for most of training. Falling means episodes terminate early due to coverage > 95% (good) or crashes (bad).
  • loss/value_loss decreasing over time. Typical range: starts at ~5, ends at ~0.5.
  • loss/policy_loss small and stable (~0.01-0.05). Sharp spikes mean the policy is updating too aggressively.
  • train/entropy_loss going from ~−2.0 (uniform action choice) to ~−1.4 (focused) over 1M steps. If it crashes toward 0, the policy collapsed.
  • train/explained_variance approaching 1 (e.g., 0.95-0.99 by end). If it stays low, the value function isn’t learning.
  • time/fps stable. Drops indicate I/O contention or memory pressure.

What regressions look like

Reward curve plateaus low

Coverage in the 17-19% range across 1M steps. Classic symptom of the physical-ceiling problem before EXP-7. The fix isn’t algorithmic — it’s the action space.

Entropy collapse

Entropy loss heads to 0 within 50k steps. Policy picks one action and never explores. Cure: raise ent_coef (we settled on 0.05).

Value function divergence

explained_variance heads to negative numbers. Means the value function is anti-predicting returns. Usually a bug in reward computation (e.g., a NaN slipping through).

TensorBoard missing ep_rew_mean

If the ep_rew_mean curve is just absent, the env isn’t wrapped in Monitor. See first PPO run for the canonical recipe.

Per-experiment comparison

When running comparison experiments, the most useful view is overlaying ep_rew_mean for baseline vs experiment. Subtle improvements (or regressions) within noise need ≥3 seeds per condition to be meaningful — single-seed comparisons mislead more often than they inform.

Where to go next

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR