claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

04 — RND (Random Network Distillation) for exploration: nothing

RND adds a learned novelty bonus on top of PPO: a predictor network mirrors a frozen target, and the mismatch rewards exploring unfamiliar ground.

stablerl-labupdated 2026-05-11T00:00:00.000ZClaudeDroneRLDevLog

What this is: we added a modern “novelty bonus” to training — RND. It uses two networks: a target that stays random and frozen, and a predictor that learns to imitate the target on whatever the agent has already seen. On familiar observations the predictor matches well; on unfamiliar ones it predicts poorly, and that gap becomes an extra “go explore” incentive. Why: the frontier-based incentive from the previous experiment may have been too hand-tuned. RND instead learns what counts as “new” on its own. It is a standard choice in 2024-2025.

Date: 8 May 2026 Machine: D2


1. What we wanted

RND idea (Burda et al. 2018): two networks.

  • Target — random and frozen.
  • Predictor — learns to imitate the target on observations the agent has already encountered.

The difference between their outputs on an unfamiliar observation acts as a “haven’t seen this before” signal that nudges the agent to explore.

Hypothesis: on our task, coverage should rise from the 19.4% baseline (NEXT-1) toward 60-75%.

2. What we did

Installed rllte-core 1.0.1. It provides RND, but no ready-made integration with Stable-Baselines3 (our main library).

Wrote our own wrapper, VecRNDWrapper (about 80 lines). On each step it pulls the visited-area information out of the observations, runs it through the RND computation to produce an internal novelty signal, and blends a small amount of that signal into the regular task reward.

Three bugs surfaced along the way during integration (wrapper attributes, tensor dimensions, gradient flow). All were fixed. The wrapper is reusable for future exploration algorithms (ICM, NGU, RIDE) because they share the same API.

3. What we got

Training

  • ep_rew_mean rose from 167 to 678 (vs NEXT-1: 709, EXP-2 frontier: 869) — lower than both.
  • The internal novelty signal grew from roughly 0.6 to 1.27 — alive, and it did not collapse.
  • Throughput dropped to 1651 fps from 2700 fps without RND — about 40% overhead.

Evaluation

Metric EXP-3 RND EXP-2 frontier NEXT-1 baseline
coverage stoch @1k 17.8% 19.0% 19.4%

RND came out slightly worse than the baseline. That is the third null result in a row, with the curve flattening around 19%.

4. Why it didn’t work

Because the task already rewards every newly visited cell, the drone is constantly at the edge of explored territory — almost every step encounters fresh ground. RND therefore never finds a genuinely unexplored spot; nearly everything looks “unseen before” to roughly the same degree.

Effects working against us:

  1. Extra noise for advantage estimation. PPO assumes the reward signal is stationary. The novelty signal, by contrast, shifts as the predictor keeps learning, which lowers the signal-to-noise ratio for the policy gradient.
  2. No gradient toward the truly unexplored. The drone sits permanently at the boundary of known territory, so all steps look about equally novel.

Third confirmation of the ~25% physical ceiling

Three approaches (CNN with domain randomization, frontier-based incentive, and RND) all flattened in the 17-19% range. This points to a structural limit:

  • The map holds roughly 3700 free cells.
  • The drone moves one cell per step.
  • With an episode length of 1000 steps, the realistic maximum is about 25% (around 1000 cells once you account for inefficiency).

No amount of reward tuning will move that ceiling when the agent has 1000 steps and moves a single cell at a time.

5. What’s next

  • TASK-RL-EXP-4 — retrain with much longer episodes (3000 steps). If the ceiling is tied to the horizon, this should lift it. (Spoiler: also a null result; see dev-log/06.)
  • TASK-RL-NEXT-3 — implement the classic lawnmower heuristic for an honest comparison. (Spoiler: it turned into a sharp wake-up call.)
  • Multi-cell forward motion — still in the backlog, but a promising candidate. (Spoiler: it turned out to be the key; see dev-log/07.)

What we’ll reuse

VecRNDWrapper is written so that RND can be swapped for any other internal-reward algorithm (ICM, RIDE, NGU, RE3) with a single change. We’re keeping it.

Related files

  • training/rnd_wrapper.py (new)
  • training/callbacks.py — RNDStatsCallback for logging
  • training/configs/ppo_rnd.yaml
  • experiments/TASK-RL-EXP-3/

Glossary

  • RND (Random Network Distillation) — an exploration algorithm. Burda et al. 2018, “Exploration by Random Network Distillation.”
  • Intrinsic reward — an internal reward the agent gives itself, separate from the task, meant to encourage exploration.
  • Extrinsic reward — the external reward that comes from the task itself.
  • Predictor / Target — the two RND networks. The target is frozen and random; the predictor learns to imitate it.
  • rllte-core — a library (RLE Foundation) with modern exploration algorithms.
  • VecEnvWrapper — a wrapper around a vectorized environment in SB3; the integration point for our code.
  • fps — frames per second; training throughput.
  • Frontier — the boundary between known and unknown territory.
  • Physical ceiling — the structural limit that comes from the geometry of the task itself.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR