What this is: the drone had a set of eight moves, one being “scan” — rotating the lidar servo. We removed it and let the drone learn with the remaining seven. Why: earlier analysis suggested scan contributed nothing. The drone was spending a notable share of its time on a move that covered no new cells. If it is just noise, removing it would simplify the future hierarchical design.
Date: 9 May 2026 Machine: D2
1. What we wanted to check
Context: while analysing our best baseline run we noticed:
- The scan move covers no new cells — its contribution to coverage was essentially nothing.
- Yet the drone chose it for a meaningful fraction of every long episode.
- It earned a small positive reward each time, a bonus tied to taking a rangefinder reading.
Suspicion: scan is noise. The drone had learned a workaround — collecting that small bonus in bad spots instead of accepting the usual penalty for moving.
First idea: if we remove scan, the model should redirect that freed-up budget to useful moves, keeping coverage the same or improving it. Second idea: scan indirectly helps — as a forced pause and a sensor refresh — so without it coverage would fall noticeably.
2. What we did
We built a variant of the environment that is a copy of the main one but with the scan move taken out:
- The drone now chooses from a discrete set of seven moves instead of eight.
- The servo is frozen pointing straight ahead and never rotates.
- Everything else is identical.
A quick one-minute sanity run trained without crashing.
Full training over roughly a million steps (about six minutes):
- Reward per episode was clearly higher than the baseline — better during training.
- Throughput matched the baseline.
- The model settled into confident, decisive action choices.
We then evaluated on the same five unseen maps used for the baseline.
3. What we saw
Average coverage — a mild, uniform regression
| Step limit | No scan | With scan (baseline) | Δ |
|---|---|---|---|
| 1000 | 54.80% | 56.10% | −1.30 pp |
| 3000 | 83.15% | 86.18% | −3.03 pp |
| 5000 | 92.26% | 93.75% | −1.49 pp |
The “scan is just noise” idea did not hold up: removing it made coverage worse at every limit. The “scan indirectly helps” idea held up in part — the regression was real and consistent, though never dramatic (always under a couple of points).
Per map — the model never reaches 95%
| Map | No scan | With scan (baseline) | Δ | Episode length |
|---|---|---|---|---|
| map_00 | 91.75% | 94.10% | −2.35 | 5000 (always max) |
| map_01 | 92.80% | 95.05% | −2.25 | 5000 |
| map_02 | 91.21% | 93.38% | −2.17 | 5000 |
| map_03 | 91.55% | 91.76% | −0.21 | 5000 |
| map_04 | 94.01% | 94.48% | −0.47 | 5000 |
Qualitative difference: the baseline sometimes ends episodes early, once it reaches 95% coverage. Without scan the model never reaches 95%, so the episode always runs all the way to the limit.
The “trains better, evaluates worse” pattern — a fifth time in a row
In training the model earns more reward; on new maps it does slightly worse. This is now the fifth run in a row showing the same shape: stronger during training, weaker on held-out evaluation.
4. What this means
Headline: scan is not noise. It has a hidden role:
- A forced pause. When the drone is boxed into a dead end and any movement is penalised, scan lets it wait in place while still collecting a small reward. Without it, the model is forced to spin or bump into a wall, both of which are penalised.
- An observation refresh. The rangefinder reading depends on where the servo points, so a scan changes what the drone sees next — an implicit source of variety that aids exploration.
- Possible overuse of the long-forward move. With the scan budget gone, that time may shift into driving forward until an obstacle, which could cause extra collisions. We did not verify this; it would need the action-distribution tooling adapted to seven moves.
For the future hierarchical design:
- The low-level behaviour should include a “pause and wait” mechanism, an analogue of scan.
- The size of the move set is best decided after a proper review of the literature.
The pattern, five times over:
Five experiments in a row tell the same story: training improves while evaluation stays flat. That is a strong sign the current training-map pool may simply be too small to generalise. A plausible next step is to expand the pool considerably and retrain the baseline — a cheaper move than a structural redesign.
5. What’s next
- The hierarchical direction is the top priority, preceded by a thorough review of the relevant literature.
- Expand the training-map pool substantially and retrain the baseline. If evaluation improves, the ceiling was in the data, not the architecture — a cheap way to settle the question.
- Optionally, analyse how the seven moves are distributed in use, to check whether the long-forward move is being overused.
Glossary
- scan move — rotating the servo that carries the rangefinder. It does not move the drone but grants a small bonus for taking a reading.
- Rangefinder — the distance sensor; in our model, a value in the drone’s observations.
- Servo — a rotating mechanism, with an angle between 0 and 180 degrees.
- Move set — the fixed list of discrete moves the drone can choose from. Ours had eight: forward, back, left, right, turn one way, turn the other, scan, and a long forward move that runs until an obstacle.
- Long forward move — driving forward across several cells until blocked. The main growth driver in our best baseline.
- Sanity run — a short run to confirm the code does not crash.
- Aggregate — the mean across all maps and episodes.
- Out-of-distribution — data the model did not see during training.