TensorBoard is the main diagnostic tool during training. It shows you in near-real-time whether the policy is learning, whether the run is degenerate, and whether something is broken. This page covers what to watch for.
Healthy training curves
For a well-tuned PPO run on our coverage task, you should see:
ep_rew_meanclimbing from ~160 (random) to ~2000 (converged) over 1M steps. Some noise is normal; sudden drops are a red flag.ep_len_meanstable at 1000 (max episode length) for most of training. Falling means episodes terminate early due to coverage > 95% (good) or crashes (bad).loss/value_lossdecreasing over time. Typical range: starts at ~5, ends at ~0.5.loss/policy_losssmall and stable (~0.01-0.05). Sharp spikes mean the policy is updating too aggressively.train/entropy_lossgoing from ~−2.0 (uniform action choice) to ~−1.4 (focused) over 1M steps. If it crashes toward 0, the policy collapsed.train/explained_varianceapproaching 1 (e.g., 0.95-0.99 by end). If it stays low, the value function isn’t learning.time/fpsstable. Drops indicate I/O contention or memory pressure.
What regressions look like
Reward curve plateaus low
Coverage in the 17-19% range across 1M steps. Classic symptom of the physical-ceiling problem before EXP-7. The fix isn’t algorithmic — it’s the action space.
Entropy collapse
Entropy loss heads to 0 within 50k steps. Policy picks one action and never explores. Cure: raise ent_coef (we settled on 0.05).
Value function divergence
explained_variance heads to negative numbers. Means the value function is anti-predicting returns. Usually a bug in reward computation (e.g., a NaN slipping through).
TensorBoard missing ep_rew_mean
If the ep_rew_mean curve is just absent, the env isn’t wrapped in Monitor. See first PPO run for the canonical recipe.
Per-experiment comparison
When running comparison experiments, the most useful view is overlaying ep_rew_mean for baseline vs experiment. Subtle improvements (or regressions) within noise need ≥3 seeds per condition to be meaningful — single-seed comparisons mislead more often than they inform.
Where to go next
- First PPO run (dev-log 02) — pipeline bring-up including TensorBoard wiring
- PPO hyperparameters — what the config tunes
- Training metrics hub — sibling pages