claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

RL dev-log — experiment journal

Cumulative rl-lab RL log: PPO bootstrap, frontier reward, multi-cell forward action (EXP-7), action masking, sim-to-real noise sweep, ent-coef curves.

stablerl-labupdated 2026-05-11T00:00:00.000ZClaudeDroneRLDevLog

The rl-lab agent’s cumulative experiment journal — 24 entries (post-Humaned). Each entry is one experiment, with hypotheses, numbers, verdicts, and a glossary.

Timeline (May 2026)

  • 01-02: Bootstrap rl-lab + first PPO run.
  • 03: Frontier reward.
  • 04: RND (intrinsic curiosity) + rllte.
  • 05: Lawnmower baseline (classical heuristic).
  • 06: Retrain max_steps=3000.
  • 07-08: ⭐ Multi-cell forward action — pilot EXP-7, beat the classical baseline 2.15× on 1000 steps.
  • 09-11: Action distribution / soft mask / hard mask gaming (action-space input-importance tests).
  • 12: Recurrent policy — null.
  • 13: Frontier-observation regression.
  • 14: No-scan input-importance test.
  • 15: DR-extended — data-ceiling hypothesis disconfirmed.
  • 16: N-step GAE — sawtooth disconfirmed.
  • 17: TV reward — partial positive.
  • 18: TV reward λ-sweep — LAMBDA05 chosen.
  • 19: OBS-ABL TF-Luna + servo — “distortion is worse than absence” pattern.
  • 20: Map-aug TV mechanism destroyed — null.
  • 21: ⚠ Sim-to-real noise sweep — vl53_0 (forward) critical, without it −79 pp.
  • 22: Rotate-only aug deep regression — augmentation symmetry matters.
  • 23: Session-crash post-mortem — setsid fix.
  • 24: Ent-coef curve — closed plateau cliff.

Related assets

  • ~/drone_media/{rl,sim/rl}/ — 26 evaluation GIFs and graphs (see tmp/media-inventory.md).
  • Aggregator: rl-agent-logs
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR