What this is: we added an extra map to the drone’s observations describing how far the nearest unexplored cell is from each point. A logical “where to head” hint. Why: the previous attempt at adding memory produced no measurable change. We decided to give a more direct hint: not “where I have been,” but “where I should go.” We expected a simple, qualitatively new improvement.
Date: 8 May 2026 Machine: D2
1. What we wanted to test
After the no-change result from memory in the previous experiment, we shifted to informational changes to what the drone observes.
The frontier-map idea:
- The drone already tracks a map of where it has been (visited versus not visited).
- We add a same-sized map encoding, for every point, how far it is to the nearest unexplored area.
- This is a qualitatively different signal: not “where I have been” (memory), but “where I should go” (a goal).
Implementation: a standard distance-transform routine converts the visited map into a distance-to-unexplored map.
What we expected:
- Good case: coverage at the 5,000-step horizon rises from 93.4% toward 96-98%, closing the gap to the classical lawnmower heuristic (98.7%).
- Bad case: the extra map is noisy, or the distance hint is not actionable, leading to a regression.
2. What we did
We built a new environment that extends the drone’s observations with the distance-to-unexplored map.
We added a new feature extractor running two convolutional branches in parallel:
- One branch processes the visited map, as before.
- The second branch processes the distance-to-unexplored map.
- Their outputs are combined with the rangefinder distances and the servo angle into a single feature vector.
Training ran for one million steps across 16 parallel environments in roughly 6.8 minutes:
- Training reward came out about 4% above baseline (better during training).
- Throughput was around 2,491 steps per second.
3. What we saw
The regression grows with episode length
| Step limit | Experiment | Baseline EXP-7 | Difference |
|---|---|---|---|
| 1000 | 54.67% | 56.10% | −1.43 (within noise) |
| 3000 | 79.48% | 86.18% | −6.70 points |
| 5000 | 86.20% | 93.75% | −7.55 points ⚠ |
On longer episodes the gap widens. And the drone never reaches 95% — every run uses the full step budget, while the baseline sometimes finishes earlier.
Versus the classical lawnmower: the gap widened from about −5 points at baseline to roughly −12.5 points. The extra map is not neutral — it actively hurts.
Across all five maps — a uniform decline
Every map lost between 5 and 10 points. This is not a single outlier; the effect is systemic.
4. What it means
The expectation that “more information in the observations is better” did not hold here.
Why the extra map hurts:
- “Somewhere far away” is not actionable. The drone may see that there is unexplored area dozens of cells to one side, but it acts locally — a step forward, or forward until it would collide. A global signal does not translate into a local action.
- Two convolutional branches on the same compute budget learn worse. The learning signal is split between the useful branch (visited) and the distracting one (distance-to-unexplored), leaving both weaker than a single focused branch.
- It disrupts the learned traversal pattern. Earlier analysis found the drone had taught itself a snake-like sweep — turn, turn, then a long run forward. The “go to the far corner” hint pulls it off that efficient pattern.
A methodological lesson:
- In the published literature, this kind of frontier information is typically used as part of the reward — a bonus for moving toward unexplored area — rather than as an observation. As an observation in our setup it does not work; an earlier reward-based attempt also showed no change, but that used an older single-cell action set. It may behave differently with the current multi-cell movement.
This is the fourth informational or architectural change in a row that failed to help:
- Soft action constraints — no change
- Hard action constraints — collapsed, with the model gaming the metric
- Recurrent memory — no change
- This distance-to-unexplored map — a regression of about 7.5 points
The previous best model (EXP-7) remains the production model.
5. What’s next
After four informational and architectural dead ends, we plan to pivot toward structural or environmental changes:
- Hierarchy — the most promising direction: a high level that picks a direction and a low level that handles the stepping. A structural prior, without the gaming risk and without piling on redundant information.
- A simpler environment variant that removes the scan step — cheap to try.
- Frontier information as a reward rather than an observation — worth re-evaluating with the current multi-cell movement.
Related files
- The frontier-augmented environment
- The feature extractor (with its two-branch variant)
- The matching training and evaluation scripts
- The associated experiment directory
Glossary
- Observation — what the model sees each step: rangefinder distances, servo angle, and the map of where it has been.
- Frontier — the boundary between known and unknown space. A standard term in coverage path planning.
- Distance-to-unexplored map — for every cell, the distance to the nearest unexplored cell, normalized to a 0-to-1 range.
- Distance transform — a standard computer-vision algorithm that gives every point its distance to the nearest point of another class.
- CNN — a convolutional neural network; it processes maps like images.
- Extractor — the part of the policy network that turns raw observations into features before the main decision layer.
- Boustrophedon — a snake-shaped, back-and-forth traversal.
- Lawnmower — a classical, non-learning coverage heuristic. Our reference reaches 98.7% coverage at 5,000 steps.
- EXP-7 — the current best baseline model.