claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

24 — Tuning the exploration setting: finding the stable range for indoor coverage

Swept the policy's exploration setting and found a stable sweet-spot range, with a sharp drop-off when pushed too far.

stablerl-labupdated 2026-05-11T00:00:00.000ZClaudeDroneRLDevLog

What this is: I swept the policy’s exploration setting across several trainings and mapped how coverage responds. The response is asymmetric: a wide stable range where results barely change, then a sharp drop-off once the setting goes too high. Why it’s here: a production decision — we keep the current best model as the base. Plus a bonus: an explanation of why four morning trainings died (not a “session crash” but a container memory limit).

Date: 2026-05-10


Glossary

  • Exploration setting — a knob in the PPO training loop that controls how much the policy is pushed to try new actions versus exploiting the strategy it has already learned. Turn it up and the agent experiments more; turn it down and it sticks to what it knows.
  • PPO — RL algorithm (Proximal Policy Optimization). Our default.
  • Coverage — % of map cells the drone visited. Main metric.
  • Lawnmower — classical “mower” coverage algorithm (no RL). Gives 98.7% at 5k steps. The reference we’re trying to beat by 0.5%.
  • OOM (Out of Memory) — OS kills a process when out of RAM.
  • cgroup — isolated “container” for processes with RAM and CPU limits.

1. What I wanted to check

We already had two reference points: a moderate exploration level that gave our best coverage so far, and a very high level that was a catastrophe (about −18.5 pp). What we did not know was the shape between them — was the good region a narrow point or a broad band? So I wanted to fill in the gaps to:

  1. Understand where coverage starts to fall off.
  2. Confirm the current best really is the best.
  3. Possibly find an even better setting (if the optimum turned out elsewhere).

2. What I tried

I ran several PPO trainings, 1M steps each, varying only the exploration setting. Each took about 8 minutes on the GPU (RTX 5070).

Morning blocker: I launched four in parallel — all died within 20 seconds. I initially thought it was a “Claude session crash” from a wrong setsid — it turned out to be OOM from the cgroup. Claude Code runs in a “container” with a 16 GB RAM limit. Four trainings together eat ~20 GB — it doesn’t fit, so the kernel kills them.

Fix (~25 minutes): in scripts/parallel_trains.sh I added a memory check before launch. If four won’t fit, the script proposes batching into pairs.

Relaunch: batch 1 (two trainings in parallel, 8 minutes), batch 2 (two trainings run sequentially, by Aleks’s call, for safety, 22 minutes).

3. What I got

The picture came out clearly asymmetric.

There is a stable range of exploration settings where coverage barely moves — two different settings inside that band landed within a fraction of a point of each other at 3k steps. That tells us the model is robust to small changes in this knob, which is reassuring: we are not balanced on a knife edge.

Below that range, dialling exploration too low also hurt a little (coverage dipped about a point). My intuition had been “less exploration is better once the policy is trained” — that was wrong. A healthy minimum of exploration is needed.

Above the stable range, the result is not a gradual decline but a sharp drop-off — coverage fell by roughly ten points across a small change in the setting. I did not expect such an abrupt transition. It behaves like a phase change in PPO training: past a certain point the exploration pressure starts to dominate the learning signal and the strategy disintegrates. Push the setting all the way up and you get the full catastrophe we saw earlier.

One nuance worth keeping: a slightly higher (but still in-range) setting did a little better on corridor maps (map_03, ~94.4% vs ~93.5%). A touch more exploration seems to help in difficult geometry, so it’s a reasonable backup for corridor-heavy missions.

Decision

  • Production base: the current best setting, in the middle of the stable range.
  • Backup: the slightly higher in-range setting for corridor domains.
  • Avoid: anything above the stable range (the drop-off zone) or below the healthy minimum (the over-exploitation valley).

Mission progress

Goal: beat lawnmower (98.7%) by at least 0.5%. Currently:

  • The old base sat at 94.1% → gap −4.6 pp
  • The new best is at 94.9% → gap −3.8 pp

So the new base moved 0.8 pp ahead of the old one. We still need to close 3.8 pp to lawnmower, plus 0.5 pp for the mission goal. The next direction is map augmentation, which the literature suggests could add several points.

4. Sources

  • PPO original paper (Schulman 2017, arxiv:1707.06347) — why an exploration term helps training.
  • rl-baselines3-zoo — typical exploration-setting ranges for discrete action spaces like ours; our stable range lines up with the common guidance.
  • systemd cgroup v2 docs (man systemd.resource-control) — about memory.max and why setsid doesn’t unbind the child process from the cgroup. This explained why the morning runs kept failing.

Bonus: cgroup OOM breakdown (for future sessions)

Why it matters: this morning I lost ~30 minutes on a wrong diagnosis (“setsid didn’t work”) instead of the real one (“cgroup OOM”). To avoid repeats — this section.

OOM symptoms (distinguishing from a real crash):

  • All trainings die simultaneously 15-25 seconds after start.
  • The last log line is Logging to /home/.../runs/... (the moment PPO spawns 16 worker processes for parallel rollout).
  • No model.zip.
  • The bash session is still alive (in a real crash, bash would die too).

How to check in 5 seconds:

journalctl --since "1 hour ago" | grep -iE "oom|killed.*python"

If it shows “Memory cgroup out of memory: Killed process X (python)” — it’s OOM.

Claude session numbers:

  • memory.max: 16 GB (hard limit)
  • 1 training takes: ~5 GB (1 GB main + 16 envs × 256 MB)
  • 4 in parallel: ~20 GB → doesn’t fit → OOM

Safe configurations:

  • 1 training sequential: ~5 GB peak — always OK
  • 2 parallel: ~10 GB peak — OK with margin
  • 3+ parallel: NO, batch via parallel_trains.sh launch-batched <specs> 2
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR