What this is: we took our two coverage models apart (an earlier model that was a notable step forward, and a later one that tried to improve on it). We counted how often each action is chosen, how many cells it covers, and what sequence of actions emerges during a run. Why: the earlier model was a notable step forward and the later one regressed slightly. The qualitative guesses were already in hand; we needed numbers to understand why.
Date: 8 May 2026 Machine: D2
1. What we wanted to check
The drone can choose one of eight actions:
- a0: one cell forward
- a1: one cell back
- a2: one cell left (strafe)
- a3: one cell right (strafe)
- a4: rotate by a small positive angle
- a5: rotate by a small negative angle
- a6: scan (servo rotation; does not move the drone)
- a7: forward-until-collision (advance toward an obstacle — many cells in a single call)
We started from three questions. First, whether the stronger model leans heavily on the long-advance action. Second, whether the weaker model regressed because it favors rotate and scan instead of that long advance. Third, whether the stronger model had learned a snake-like sweep — turn, long move, turn, long move.
2. What we did
We wrote an evaluation script that, for each step, records which action was chosen, how many new cells it produced (the change in visited count before and after), and what the previous action was, so we could build a transition matrix.
We ran four variants: the stronger model at a short step budget (its training setting) and at a long step budget (its production setting), and the weaker model at the same two budgets. Each variant used five maps and five episodes per map, twenty-five episodes per variant and one hundred episodes in total, taking roughly six minutes on the GPU.
3. What we saw
Action frequencies
| Action | Stronger @short | Stronger @long | Weaker @short | Weaker @long |
|---|---|---|---|---|
| 1-cell moves (a0-a3) | 17% | 24% | 14% | 27% |
| rotates + scan (a4-a6) | 70.4% | 65.3% | 73.2% | 64.2% |
| action 7 | 12.7% | 10.9% | 12.2% | 8.7% |
We found that the long-advance action is chosen only 8-13% of the time in every variant, far less than we had assumed.
Action contribution to coverage — the main observation
| Group | Stronger @short | Stronger @long |
|---|---|---|
| 1-cell moves (4 actions) | 2.1% | 2.3% |
| rotates + scan (3 actions) | 0% | 0% |
| action 7 (1 action) | 98.0% | 97.7% |
The result is striking: action 7 produces about 98% of all coverage while being chosen only around 10% of the time. Every other action is essentially positioning for the next action 7 call. Scan, rotate, and strafe on their own cover almost nothing.
In other words, dominance is not about how often the action fires — it is about how much it contributes per call. That is a sharper and more interesting finding than the simple frequency story we started with.
Action 7 efficiency — cells per call
| Stronger @short | Stronger @long | Weaker @long | |
|---|---|---|---|
| Number of calls | 3182 | 13142 | 10792 |
| Mean cells per call | 16.14 | 6.51 | 7.86 |
| Median | 12 | 2 | 4 |
| Max | 61 | 61 | 61 |
On the long episodes, cells per call drops from about sixteen to six-to-eight, because most of the map has already been visited and the long advance ends up traversing already-covered cells with no gain.
There is a small paradox in the weaker model: on the long budget it executes the long advance slightly more efficiently (7.86 versus 6.51 cells per call), yet it chooses that action less often (8.7% versus 10.9%). Its regression is therefore explained by lower frequency, not by weaker execution per call.
Transition matrix — the snake pattern
Stronger model on the short budget:
- rotate-positive followed by rotate-positive: 65.8% (several rotates in a row to align)
- rotate-negative followed by rotate-negative: 68.0%
- action 7 followed by a rotate (either direction): about 79% of the time a long move is followed by a turn
This confirms a snake-like traversal. The policy learned a clear rhythm: align with a run of rotates, make a long move, align again, make another long move — exactly the classical lawnmower pattern.
Weaker model on the long budget:
- rotate-positive followed by rotate-positive: 52.5% (versus 65.8%) — weaker
- action 7 followed by a rotate: 64% (versus 79%) — weaker
The weaker model produces less structured behavior, more often following a long move with scan, strafe, or a step back instead of the turn that would set up the next pass.
4. What this means
The headline is that action 7 is the limiting factor for the whole model: a single action produces 98% of the result.
For strategy this suggests a few directions. Raising the frequency of the long-advance action well above its current level could yield a large coverage gain, for example by masking out unhelpful actions or by adding a higher-level option that sweeps until collision. A policy with memory of where it has already been might raise the cells produced per call, because it could avoid driving straight ahead into already-visited space. And scan, which contributes nothing to coverage, is a candidate to remove entirely.
The fact that the policy independently discovered the classical snake pattern is encouraging: it means the learning is working in a structured way rather than by accident.
This also looks like a new empirical result — we have not seen comparable numerical action-distribution analyses for coverage-path-planning RL in the literature. It is a possible paper candidate if the finding holds up on other models.
5. What’s next
We opened five follow-up directions:
- action masking — promising and cheap.
- removing the scan action — a simple input-importance test.
- a recurrent policy with memory.
- a hierarchical option that pairs alignment with a sweep, mirroring the learned pattern.
- exposing frontier information in the observation, as a signal for where to head next.
(As an aside: by the time this write-up was finished, all five had turned out to be null results or regressions. But they only became possible and well-justified through this analysis.)
Related files
eval_action_dist— action counter plus transition matrixplot_action_dist— plots- per-variant action-distribution data files
- action-distribution and transition grid figures
Glossary
- Stronger model — our earlier baseline with the multi-cell action; a notable step forward at the time.
- Weaker model — a later attempt to improve on it that regressed.
- action 7 (forward-until-collision) — the multi-cell “advance to obstacle” action, the main productive action.
- Coverage contribution per action — how many new cells a given action covers, summed across all episodes.
- Cells per call — how many new cells one action 7 call produces on average.
- Transition matrix — how many times each “action X then action Y” sequence occurred.
- Boustrophedon — a snake-shaped traversal pattern, a classical coverage approach.
- Lawnmower — a classical (non-RL) lawnmower-shaped coverage heuristic.
- Stochastic evaluation — the model picks actions by sampling from the probability distribution rather than always taking the most likely one.