claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

13 — 'Distance to unexplored' map in observations: worse by 7.5 points

Adding a 'distance to unexplored' map to drone observations hurt indoor coverage by 7.5 points. Why richer input did not help.

stablerl-labupdated 2026-05-11T00:00:00.000ZClaudeDroneRLDevLog

What this is: we added an extra map to the drone’s observations describing how far the nearest unexplored cell is from each point. A logical “where to head” hint. Why: the previous attempt at adding memory produced no measurable change. We decided to give a more direct hint: not “where I have been,” but “where I should go.” We expected a simple, qualitatively new improvement.

Date: 8 May 2026 Machine: D2


1. What we wanted to test

After the no-change result from memory in the previous experiment, we shifted to informational changes to what the drone observes.

The frontier-map idea:

  • The drone already tracks a map of where it has been (visited versus not visited).
  • We add a same-sized map encoding, for every point, how far it is to the nearest unexplored area.
  • This is a qualitatively different signal: not “where I have been” (memory), but “where I should go” (a goal).

Implementation: a standard distance-transform routine converts the visited map into a distance-to-unexplored map.

What we expected:

  • Good case: coverage at the 5,000-step horizon rises from 93.4% toward 96-98%, closing the gap to the classical lawnmower heuristic (98.7%).
  • Bad case: the extra map is noisy, or the distance hint is not actionable, leading to a regression.

2. What we did

We built a new environment that extends the drone’s observations with the distance-to-unexplored map.

We added a new feature extractor running two convolutional branches in parallel:

  • One branch processes the visited map, as before.
  • The second branch processes the distance-to-unexplored map.
  • Their outputs are combined with the rangefinder distances and the servo angle into a single feature vector.

Training ran for one million steps across 16 parallel environments in roughly 6.8 minutes:

  • Training reward came out about 4% above baseline (better during training).
  • Throughput was around 2,491 steps per second.

3. What we saw

The regression grows with episode length

Step limit Experiment Baseline EXP-7 Difference
1000 54.67% 56.10% −1.43 (within noise)
3000 79.48% 86.18% −6.70 points
5000 86.20% 93.75% −7.55 points ⚠

On longer episodes the gap widens. And the drone never reaches 95% — every run uses the full step budget, while the baseline sometimes finishes earlier.

Versus the classical lawnmower: the gap widened from about −5 points at baseline to roughly −12.5 points. The extra map is not neutral — it actively hurts.

Across all five maps — a uniform decline

Every map lost between 5 and 10 points. This is not a single outlier; the effect is systemic.

4. What it means

The expectation that “more information in the observations is better” did not hold here.

Why the extra map hurts:

  1. “Somewhere far away” is not actionable. The drone may see that there is unexplored area dozens of cells to one side, but it acts locally — a step forward, or forward until it would collide. A global signal does not translate into a local action.
  2. Two convolutional branches on the same compute budget learn worse. The learning signal is split between the useful branch (visited) and the distracting one (distance-to-unexplored), leaving both weaker than a single focused branch.
  3. It disrupts the learned traversal pattern. Earlier analysis found the drone had taught itself a snake-like sweep — turn, turn, then a long run forward. The “go to the far corner” hint pulls it off that efficient pattern.

A methodological lesson:

  • In the published literature, this kind of frontier information is typically used as part of the reward — a bonus for moving toward unexplored area — rather than as an observation. As an observation in our setup it does not work; an earlier reward-based attempt also showed no change, but that used an older single-cell action set. It may behave differently with the current multi-cell movement.

This is the fourth informational or architectural change in a row that failed to help:

  • Soft action constraints — no change
  • Hard action constraints — collapsed, with the model gaming the metric
  • Recurrent memory — no change
  • This distance-to-unexplored map — a regression of about 7.5 points

The previous best model (EXP-7) remains the production model.

5. What’s next

After four informational and architectural dead ends, we plan to pivot toward structural or environmental changes:

  • Hierarchy — the most promising direction: a high level that picks a direction and a low level that handles the stepping. A structural prior, without the gaming risk and without piling on redundant information.
  • A simpler environment variant that removes the scan step — cheap to try.
  • Frontier information as a reward rather than an observation — worth re-evaluating with the current multi-cell movement.

Related files

  • The frontier-augmented environment
  • The feature extractor (with its two-branch variant)
  • The matching training and evaluation scripts
  • The associated experiment directory

Glossary

  • Observation — what the model sees each step: rangefinder distances, servo angle, and the map of where it has been.
  • Frontier — the boundary between known and unknown space. A standard term in coverage path planning.
  • Distance-to-unexplored map — for every cell, the distance to the nearest unexplored cell, normalized to a 0-to-1 range.
  • Distance transform — a standard computer-vision algorithm that gives every point its distance to the nearest point of another class.
  • CNN — a convolutional neural network; it processes maps like images.
  • Extractor — the part of the policy network that turns raw observations into features before the main decision layer.
  • Boustrophedon — a snake-shaped, back-and-forth traversal.
  • Lawnmower — a classical, non-learning coverage heuristic. Our reference reaches 98.7% coverage at 5,000 steps.
  • EXP-7 — the current best baseline model.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR