What this is: we rewarded the drone for reaching unexplored frontier cells, and we let it explore more freely at the start of training and become more decisive toward the end. Why: the previous experiment reached 19.4% coverage in 1000 steps. We wanted to lift that into the 40-60% range using two classical ideas from the literature.
Date: 8 May 2026 Machine: D2
1. What we wanted
Idea 1: reward the drone for being close to unexplored territory, hoping it would steer toward the boundary of the explored area.
Idea 2: let the policy explore broadly early in training and gradually become more decisive, to narrow the gap between its exploratory behaviour (about 19% coverage) and its confident behaviour (about 6.5%).
2. What we did
In the environment: we added a small per-step reward tied to how close the drone was to unexplored space. It was computed with a standard computer-vision distance transform; on a 64×64 grid this took about 0.1 ms per step, negligible.
In training: a small custom routine gradually reduced the exploration pressure over the course of training.
Everything else matched the previous run.
3. What we got
Training
- Average episode reward rose from 167 to 869 (the previous run reached 709) — higher because of the new frontier bonus.
- The model’s behaviour became more decisive than in the previous run.
- No numerical instability.
Evaluation
| Metric | This run | Previous run |
|---|---|---|
| coverage (exploratory) | 19.0% | 19.4% |
| coverage (confident) | 7.2% | 6.5% |
Coverage did not grow. The frontier bonus only lifted the training reward — a roughly constant addition each step — not the actual coverage.
4. Why it didn’t work
Before running the full experiment, we measured the previous model’s behaviour and found that the drone is almost always already sitting right at the frontier during exploration. Because the base reward already gives credit for each new cell discovered, every step tends to encounter unexplored space close by.
As a result, the drone’s distance to unexplored territory was nearly always at its minimum. The frontier bonus therefore added almost the same amount every single step — effectively a constant offset rather than a signal pointing toward anywhere specific. It softened the per-step penalty but gave no real guidance about where to go.
The frontier reward would help if the drone ever found itself deep inside already-explored territory, far from any boundary. But that situation is rare while the base reward keeps the drone discovering new cells. It would mostly appear late in an episode, once more than half the area is covered — and within 1000 steps we never reach that point.
The structural limit
- There are roughly 3700 free cells.
- Covering 85% of them means about 3145 cells.
- Covering 3145 cells requires at least 3145 forward steps.
- With a 1000-step budget, the realistic maximum is around 25% coverage.
This is a limit imposed by the geometry of the task, not by training. With only 1000 steps and single-cell movement, no amount of extra reward tuning can push past it — frontier rewards, novelty-based exploration, and noisy-network exploration would all run into the same ~25% wall.
5. What’s next
- Try a novelty-based exploration method as planned. We expect it to hit the same ~25% limit.
- Raise the step budget to several thousand steps in training, so the frontier idea has a chance to matter in the later stages of an episode.
- Build a simple lawn-mower sweep baseline, so we know what classical methods can achieve.
- Consider letting the drone move several cells per action, kept in the backlog for now.
The exploration-decay routine will be reused in future experiments.
Related files
- Environment: added the frontier-distance helper and bonus to the step logic.
- Training: the exploration-decay routine.
- Experiment configuration and artifacts for this run.
Glossary
- Frontier — the boundary between known and unknown space.
- Frontier reward — a reward for being close to that boundary.
- Distance transform (EDT) — a computer-vision algorithm that gives, for every point, the distance to the nearest point of another class.
- Potential-based shaping — a way of adding helper reward that, under the Ng et al. (1999) result, provably does not change the task’s optimal solution.
- Exploratory vs. confident evaluation — running the policy with or without randomness in its choices; a large gap means the model is not yet confident.
- Novelty-based exploration — an exploration method that rewards visiting unfamiliar states.
- Noisy-network exploration — an alternative exploration method based on injecting noise into the network. Not tried yet.
- Structural limit — a ceiling that comes from the geometry of the task rather than from training.