claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

02 — First PPO run on our environment: 100k steps, reward climbs, baseline established

First full PPO training on our 2D coverage task: 100,000 steps, two early bugs fixed, reward climbs, and our first reference point captured.

stablerl-labupdated 2026-05-11T00:00:00.000ZClaudeDroneRLDevLog

What this is: the first full PPO training run on our 2D map-coverage task. 100,000 training steps, two bugs fixed in the first pass, and a final reward that climbs steadily. Why: we needed to stand up the whole pipeline — from environment to trained model — and confirm that every piece works end to end. This run is our reference point. Every later experiment is measured against it.

Date: 7 May 2026 Machine: D2 (RTX 5070)


1. What we wanted to verify

Goal: run the full cycle “environment → PPO training → evaluation → visualization” and get something working from start to finish.

Questions going in:

  • Would PPO with a MultiInputPolicy (for our dictionary-style observations) learn at least a steadily rising reward over 100k steps on a simple 64×64 room?
  • We also expected that 100k steps would be far too short for full coverage — the literature points to much longer training runs. We anticipated a rising reward but coverage well below the 95% target. This run was only ever meant to be a first reference point.

2. What we did

Environment

  • 6 VL53L0X rangefinders at angles of 0°, 60°, 120°, 180°, 240°, and 300° relative to the drone.
  • 1 TF-Luna aligned with the servo (it sweeps 0-180°, with neutral 90° pointing forward).
  • Visited map 64×64 — each cell marked 1 if visited, 0 if not.
  • Servo angle included in the observations.
  • A discrete action space of seven moves: forward, back, strafe-left, strafe-right, rotate +15°, rotate −15°, and scan.

The agent earns credit for exploring new cells and for scanning, and is penalised for collisions and for time spent, with a bonus on completing coverage. Episodes are capped at 1000 steps.

Map

A procedurally generated 64×64 layout: perimeter walls plus two horizontal and two vertical partitions, giving 3752 free cells and 344 wall cells.

Algorithm

PPO from Stable-Baselines3, with 16 parallel environments. We used the standard PPO settings for this baseline rather than tuning anything — the goal here was a working pipeline, not a tuned policy.

3. Two bugs along the way

Bug #1: the env variable didn’t take effect

Symptom: I run TRAIN_STEPS=100000 ./scripts/train.sh ..., but training overshoots to 200k and won’t stop.

Cause: train.sh was using source .env_rl, which overwrites environment variables with the values from the file. My .env_rl set a much larger step count, so it clobbered the value I passed on the command line.

Fix: I replaced source with manual parsing that only sets a variable if it isn’t already defined:

while IFS='=' read -r key value; do
  [[ -z "$key" || "$key" == \#* ]] && continue
  if [ -z "${!key:-}" ]; then
    export "$key=$value"
  fi
done < .env_rl

Now command-line parameters take precedence over the defaults in .env_rl.

Bug #2: TensorBoard showed no reward curve

Symptom: TensorBoard displayed training metrics, but there was no ep_rew_mean curve — the single most important signal.

Cause: the environment wasn’t wrapped in Monitor. SB3 relies on Monitor to detect when an episode ends and to aggregate the reward.

Fix: wrap each environment in Monitor(env) inside _make_env:

from stable_baselines3.common.monitor import Monitor

def _make_env(seed):
    def _init():
        env = Drone2DEnv(seed=seed)
        return Monitor(env)  # ← was just `return env`
    return _init

After that, the reward curve appeared.

4. Result

Training ran for 100k steps in about 42 seconds on the RTX 5070:

Step ep_rew_mean
16384 167
32768 176
65536 200
98304 219
114688 239

The reward climbed steadily across the run, and there were no NaNs and no out-of-memory issues.

Deterministic evaluation: rewards came out at [60, −5, 5, −8, 16], averaging 13.6, with coverage of 0.66%.

There was a large difference between the stochastic training reward (239) and the deterministic evaluation reward (13.6). The policy had not yet converged — under a deterministic argmax it tended to repeat a single action.

The ticket’s acceptance criterion (a rising reward in TensorBoard) was met. But 0.66% coverage is still a very long way from the 95% goal.

5. What this means for the future

What we took away:

  1. 100k steps is simply too short for PPO on a 64×64 image observation — far longer runs are the norm in the literature.
  2. For the next ticket (raising coverage past 85%), we plan to train roughly ten times longer, which still runs in only a few minutes on the RTX 5070.
  3. We may need more exploration early in training; we found that the agent settled too quickly and want to encourage broader exploration next time.
  4. It’s worth checking which feature extractor SB3 chose for our dict observation, since a convolutional extractor (NatureCNN) may be the better fit and might need to be specified explicitly.
  5. This is our very first reference point. Coverage of 0.66% is the number to beat — every later model should improve on it.

Related files

  • envs/{drone_2d_env,sensors}.py
  • maps/TASK-010/{make_simple_room.py, simple_room.png}
  • training/{train,eval}.py
  • training/configs/ppo_default.yaml
  • scripts/{train,eval,capture_results}.sh
  • experiments/TASK-010/

Glossary

  • PPO (Proximal Policy Optimization) — the reinforcement-learning algorithm we use as our main one.
  • Stable-Baselines3 (SB3) — a Python library with ready-made RL algorithm implementations.
  • MultiInputPolicy — a policy variant that accepts observations as a dict of multiple named fields. Ours combines distances, servo angle, and the visited map.
  • Dict observation space — an observation format where each field has its own name.
  • Monitor wrapper — an SB3 wrapper that tracks episode endings and emits per-episode statistics.
  • rollout — a chunk of experience collected between weight updates.
  • NatureCNN — a standard convolutional architecture from Mnih et al. 2015 (DQN on Atari).
  • CombinedExtractor — the SB3 feature extractor for dict observations; it combines multiple inputs.
  • TensorBoard (TB) — a tool for visualizing training metrics.
  • VecEnv (SubprocVecEnv) — multiple parallel environments running in separate processes.
  • Stochastic vs. deterministic evaluation — evaluation with or without sampling when selecting actions.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR