The rl-lab agent’s cumulative experiment journal — 24 entries (post-Humaned). Each entry is one experiment, with hypotheses, numbers, verdicts, and a glossary.
Timeline (May 2026)
- 01-02: Bootstrap rl-lab + first PPO run.
- 03: Frontier reward.
- 04: RND (intrinsic curiosity) + rllte.
- 05: Lawnmower baseline (classical heuristic).
- 06: Retrain
max_steps=3000. - 07-08: ⭐ Multi-cell forward action — pilot EXP-7, beat the classical baseline 2.15× on 1000 steps.
- 09-11: Action distribution / soft mask / hard mask gaming (action-space input-importance tests).
- 12: Recurrent policy — null.
- 13: Frontier-observation regression.
- 14: No-scan input-importance test.
- 15: DR-extended — data-ceiling hypothesis disconfirmed.
- 16: N-step GAE — sawtooth disconfirmed.
- 17: TV reward — partial positive.
- 18: TV reward λ-sweep —
LAMBDA05chosen. - 19: OBS-ABL TF-Luna + servo — “distortion is worse than absence” pattern.
- 20: Map-aug TV mechanism destroyed — null.
- 21: ⚠ Sim-to-real noise sweep — vl53_0 (forward) critical, without it −79 pp.
- 22: Rotate-only aug deep regression — augmentation symmetry matters.
- 23: Session-crash post-mortem —
setsidfix. - 24: Ent-coef curve — closed plateau cliff.
Related assets
~/drone_media/{rl,sim/rl}/— 26 evaluation GIFs and graphs (seetmp/media-inventory.md).- Aggregator: rl-agent-logs