What this is: the first full PPO training run on our 2D map-coverage task. 100,000 training steps, two bugs fixed in the first pass, and a final reward that climbs steadily. Why: we needed to stand up the whole pipeline — from environment to trained model — and confirm that every piece works end to end. This run is our reference point. Every later experiment is measured against it.
Date: 7 May 2026 Machine: D2 (RTX 5070)
1. What we wanted to verify
Goal: run the full cycle “environment → PPO training → evaluation → visualization” and get something working from start to finish.
Questions going in:
- Would PPO with a MultiInputPolicy (for our dictionary-style observations) learn at least a steadily rising reward over 100k steps on a simple 64×64 room?
- We also expected that 100k steps would be far too short for full coverage — the literature points to much longer training runs. We anticipated a rising reward but coverage well below the 95% target. This run was only ever meant to be a first reference point.
2. What we did
Environment
- 6 VL53L0X rangefinders at angles of 0°, 60°, 120°, 180°, 240°, and 300° relative to the drone.
- 1 TF-Luna aligned with the servo (it sweeps 0-180°, with neutral 90° pointing forward).
- Visited map 64×64 — each cell marked 1 if visited, 0 if not.
- Servo angle included in the observations.
- A discrete action space of seven moves: forward, back, strafe-left, strafe-right, rotate +15°, rotate −15°, and scan.
The agent earns credit for exploring new cells and for scanning, and is penalised for collisions and for time spent, with a bonus on completing coverage. Episodes are capped at 1000 steps.
Map
A procedurally generated 64×64 layout: perimeter walls plus two horizontal and two vertical partitions, giving 3752 free cells and 344 wall cells.
Algorithm
PPO from Stable-Baselines3, with 16 parallel environments. We used the standard PPO settings for this baseline rather than tuning anything — the goal here was a working pipeline, not a tuned policy.
3. Two bugs along the way
Bug #1: the env variable didn’t take effect
Symptom: I run TRAIN_STEPS=100000 ./scripts/train.sh ..., but training overshoots to 200k and won’t stop.
Cause: train.sh was using source .env_rl, which overwrites environment variables with the values from the file. My .env_rl set a much larger step count, so it clobbered the value I passed on the command line.
Fix: I replaced source with manual parsing that only sets a variable if it isn’t already defined:
while IFS='=' read -r key value; do
[[ -z "$key" || "$key" == \#* ]] && continue
if [ -z "${!key:-}" ]; then
export "$key=$value"
fi
done < .env_rl
Now command-line parameters take precedence over the defaults in .env_rl.
Bug #2: TensorBoard showed no reward curve
Symptom: TensorBoard displayed training metrics, but there was no ep_rew_mean curve — the single most important signal.
Cause: the environment wasn’t wrapped in Monitor. SB3 relies on Monitor to detect when an episode ends and to aggregate the reward.
Fix: wrap each environment in Monitor(env) inside _make_env:
from stable_baselines3.common.monitor import Monitor
def _make_env(seed):
def _init():
env = Drone2DEnv(seed=seed)
return Monitor(env) # ← was just `return env`
return _init
After that, the reward curve appeared.
4. Result
Training ran for 100k steps in about 42 seconds on the RTX 5070:
| Step | ep_rew_mean |
|---|---|
| 16384 | 167 |
| 32768 | 176 |
| 65536 | 200 |
| 98304 | 219 |
| 114688 | 239 |
The reward climbed steadily across the run, and there were no NaNs and no out-of-memory issues.
Deterministic evaluation: rewards came out at [60, −5, 5, −8, 16], averaging 13.6, with coverage of 0.66%.
There was a large difference between the stochastic training reward (239) and the deterministic evaluation reward (13.6). The policy had not yet converged — under a deterministic argmax it tended to repeat a single action.
The ticket’s acceptance criterion (a rising reward in TensorBoard) was met. But 0.66% coverage is still a very long way from the 95% goal.
5. What this means for the future
What we took away:
- 100k steps is simply too short for PPO on a 64×64 image observation — far longer runs are the norm in the literature.
- For the next ticket (raising coverage past 85%), we plan to train roughly ten times longer, which still runs in only a few minutes on the RTX 5070.
- We may need more exploration early in training; we found that the agent settled too quickly and want to encourage broader exploration next time.
- It’s worth checking which feature extractor SB3 chose for our dict observation, since a convolutional extractor (
NatureCNN) may be the better fit and might need to be specified explicitly. - This is our very first reference point. Coverage of 0.66% is the number to beat — every later model should improve on it.
Related files
envs/{drone_2d_env,sensors}.pymaps/TASK-010/{make_simple_room.py, simple_room.png}training/{train,eval}.pytraining/configs/ppo_default.yamlscripts/{train,eval,capture_results}.shexperiments/TASK-010/
Glossary
- PPO (Proximal Policy Optimization) — the reinforcement-learning algorithm we use as our main one.
- Stable-Baselines3 (SB3) — a Python library with ready-made RL algorithm implementations.
- MultiInputPolicy — a policy variant that accepts observations as a dict of multiple named fields. Ours combines distances, servo angle, and the visited map.
- Dict observation space — an observation format where each field has its own name.
- Monitor wrapper — an SB3 wrapper that tracks episode endings and emits per-episode statistics.
- rollout — a chunk of experience collected between weight updates.
- NatureCNN — a standard convolutional architecture from Mnih et al. 2015 (DQN on Atari).
- CombinedExtractor — the SB3 feature extractor for dict observations; it combines multiple inputs.
- TensorBoard (TB) — a tool for visualizing training metrics.
- VecEnv (SubprocVecEnv) — multiple parallel environments running in separate processes.
- Stochastic vs. deterministic evaluation — evaluation with or without sampling when selecting actions.