claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

RL Agent Work Logs — Overview

Overview of the reinforcement-learning work on autonomous 2D indoor mapping, with a chronology of dev-log stages and outcomes.

stablerl-labupdated 2026-05-11T00:00:00.000ZClaudeDroneRLAgentReports

This page gives a high-level overview of the reinforcement-learning work behind ClaudeDrone’s autonomous 2D indoor mapping. Each stage of the effort has been published in full as its own article in the RL dev-log. The work follows a familiar research arc: stand up the training infrastructure, get a first policy learning, then iterate on the reward design, the agent’s inputs, and robustness, validating each idea against measured outcomes before keeping or discarding it.

Throughout, the policy is trained with PPO and evaluated in a Gazebo-based simulator targeting deployment on ROS2 Jazzy. Progress is judged against a classical coverage baseline and against the agent’s own coverage and efficiency metrics rather than against any single internal target.

Chronology

The sequence below summarises each published entry. Early entries establish the baseline; the middle entries explore reward and input changes; the later entries focus on robustness and bringing the policy closer to real hardware.

# Stage Article
01 Bootstrap: standing up the training infrastructure from scratch 01-bootstrap
02 First PPO run 02-first-ppo-run
03 Frontier-based reward 03-frontier-reward
04 Curiosity-driven exploration via RND 04-rnd-rllte
05 Lawnmower coverage baseline 05-lawnmower-baseline
06 Retrain with a longer episode budget 06-max-steps-3000-retrain
07 ⭐ Multi-cell forward action — a notable step; this pilot beat the classical baseline by ×2.15 07-multi-cell-forward-action
08 Retrain with an extended episode budget 08-max-steps-5000-retrain
09 Action-distribution analysis 09-action-distribution-analysis
10 Soft action masking — no measurable gain 10-action-masking-soft-null
11 Hard-mask reward gaming — no measurable gain 11-hard-mask-gaming-null
12 Recurrent policy — no measurable gain 12-recurrent-policy-null
13 Frontier-observation regression 13-frontier-obs-regression
14 No-scan input-importance test dev-log 14
15 Extended domain randomization — the data-ceiling idea did not hold up 15-dr-extended-data-ceiling-disconfirmed
16 N-step advantage estimation — the expected pattern did not hold up 16-nstep-gae-saw-disconfirmed
17 Total-variation reward — partially positive result 17-tv-reward-partial-positive
18 Total-variation reward sweep — a balanced setting was selected 18-tv-reward-lambda-sweep
19 Sensor input-importance test (TF-Luna + servo can be dropped) dev-log 19
20 Map-augmentation removed the reward mechanism — no measurable gain 20-map-aug-tv-mechanism-destroyed-null
21 ⚠ Sim-to-real noise sweep — one sensor channel proved critical (−79 pp) 21-sim-to-real-noise-sweep-vl53_0-critical
22 Rotation-only augmentation — deep regression, confirming that symmetry matters 22-rotate-only-aug-deep-regression-augmentation-symmetry-matters
23 Session-crash post-mortem and fix 23-session-crash-post-mortem-setsid-fix
24 Entropy-coefficient curve — mapping the limits of further tuning 24-ent-coef-curve-closed-plateau-cliff

Outcomes at a glance

The headline result was the multi-cell forward action in stage 07, which more than doubled coverage efficiency relative to the classical baseline. Later robustness work surfaced an important finding: one sensor channel turned out to be a limiting factor under realistic noise, with a −79 pp drop when it was degraded. Several promising ideas were tested and set aside when they showed no measurable benefit, which is a normal and valuable part of the iteration.

Media

The work is illustrated by 26 evaluation animations and graphs captured during training and evaluation runs.

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR