Training metrics are how you know whether the run is healthy, whether the policy is improving, and which experiments are worth keeping. The first PPO run covers a memorable bug where the ep_rew_mean curve was missing from TensorBoard — the env wasn’t wrapped in Monitor, and we didn’t notice until we tried to diagnose why the policy “wasn’t learning” (it was, we just couldn’t see it).
TensorBoard analysis is the main page here. It covers which scalars matter, what healthy curves look like, and what regressions to watch for.
The standard scalars we monitor:
ep_rew_mean— the headline. Should climb monotonically (with noise) for most experiments.ep_len_mean— episode length. Should hit the limit (1000) on most coverage tasks; if it’s falling, episodes are terminating early due to crashes.loss/value_loss,loss/policy_loss— should decrease.train/entropy_loss— should decrease (more negative) but not to 0 (collapse).train/explained_variance— should approach 1 (typically 0.9+ for stable PPO).time/fps— throughput. Tells you if a config change tanked performance.
Contents
Auto-generated from child entries during build (update-indexes.mjs).