claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

09 — Which action does what: action 7 produces 98% of coverage while chosen 10% of the time

We took two of our coverage models apart and measured how often each action fires, how many cells it covers, and what action sequences emerge during a run.

stablerl-labupdated 2026-05-11T00:00:00.000ZClaudeDroneRLDevLog

What this is: we took our two coverage models apart (an earlier model that was a notable step forward, and a later one that tried to improve on it). We counted how often each action is chosen, how many cells it covers, and what sequence of actions emerges during a run. Why: the earlier model was a notable step forward and the later one regressed slightly. The qualitative guesses were already in hand; we needed numbers to understand why.

Date: 8 May 2026 Machine: D2


1. What we wanted to check

The drone can choose one of eight actions:

  • a0: one cell forward
  • a1: one cell back
  • a2: one cell left (strafe)
  • a3: one cell right (strafe)
  • a4: rotate by a small positive angle
  • a5: rotate by a small negative angle
  • a6: scan (servo rotation; does not move the drone)
  • a7: forward-until-collision (advance toward an obstacle — many cells in a single call)

We started from three questions. First, whether the stronger model leans heavily on the long-advance action. Second, whether the weaker model regressed because it favors rotate and scan instead of that long advance. Third, whether the stronger model had learned a snake-like sweep — turn, long move, turn, long move.

2. What we did

We wrote an evaluation script that, for each step, records which action was chosen, how many new cells it produced (the change in visited count before and after), and what the previous action was, so we could build a transition matrix.

We ran four variants: the stronger model at a short step budget (its training setting) and at a long step budget (its production setting), and the weaker model at the same two budgets. Each variant used five maps and five episodes per map, twenty-five episodes per variant and one hundred episodes in total, taking roughly six minutes on the GPU.

3. What we saw

Action frequencies

Action Stronger @short Stronger @long Weaker @short Weaker @long
1-cell moves (a0-a3) 17% 24% 14% 27%
rotates + scan (a4-a6) 70.4% 65.3% 73.2% 64.2%
action 7 12.7% 10.9% 12.2% 8.7%

We found that the long-advance action is chosen only 8-13% of the time in every variant, far less than we had assumed.

Action contribution to coverage — the main observation

Group Stronger @short Stronger @long
1-cell moves (4 actions) 2.1% 2.3%
rotates + scan (3 actions) 0% 0%
action 7 (1 action) 98.0% 97.7%

The result is striking: action 7 produces about 98% of all coverage while being chosen only around 10% of the time. Every other action is essentially positioning for the next action 7 call. Scan, rotate, and strafe on their own cover almost nothing.

In other words, dominance is not about how often the action fires — it is about how much it contributes per call. That is a sharper and more interesting finding than the simple frequency story we started with.

Action 7 efficiency — cells per call

Stronger @short Stronger @long Weaker @long
Number of calls 3182 13142 10792
Mean cells per call 16.14 6.51 7.86
Median 12 2 4
Max 61 61 61

On the long episodes, cells per call drops from about sixteen to six-to-eight, because most of the map has already been visited and the long advance ends up traversing already-covered cells with no gain.

There is a small paradox in the weaker model: on the long budget it executes the long advance slightly more efficiently (7.86 versus 6.51 cells per call), yet it chooses that action less often (8.7% versus 10.9%). Its regression is therefore explained by lower frequency, not by weaker execution per call.

Transition matrix — the snake pattern

Stronger model on the short budget:

  • rotate-positive followed by rotate-positive: 65.8% (several rotates in a row to align)
  • rotate-negative followed by rotate-negative: 68.0%
  • action 7 followed by a rotate (either direction): about 79% of the time a long move is followed by a turn

This confirms a snake-like traversal. The policy learned a clear rhythm: align with a run of rotates, make a long move, align again, make another long move — exactly the classical lawnmower pattern.

Weaker model on the long budget:

  • rotate-positive followed by rotate-positive: 52.5% (versus 65.8%) — weaker
  • action 7 followed by a rotate: 64% (versus 79%) — weaker

The weaker model produces less structured behavior, more often following a long move with scan, strafe, or a step back instead of the turn that would set up the next pass.

4. What this means

The headline is that action 7 is the limiting factor for the whole model: a single action produces 98% of the result.

For strategy this suggests a few directions. Raising the frequency of the long-advance action well above its current level could yield a large coverage gain, for example by masking out unhelpful actions or by adding a higher-level option that sweeps until collision. A policy with memory of where it has already been might raise the cells produced per call, because it could avoid driving straight ahead into already-visited space. And scan, which contributes nothing to coverage, is a candidate to remove entirely.

The fact that the policy independently discovered the classical snake pattern is encouraging: it means the learning is working in a structured way rather than by accident.

This also looks like a new empirical result — we have not seen comparable numerical action-distribution analyses for coverage-path-planning RL in the literature. It is a possible paper candidate if the finding holds up on other models.

5. What’s next

We opened five follow-up directions:

  • action masking — promising and cheap.
  • removing the scan action — a simple input-importance test.
  • a recurrent policy with memory.
  • a hierarchical option that pairs alignment with a sweep, mirroring the learned pattern.
  • exposing frontier information in the observation, as a signal for where to head next.

(As an aside: by the time this write-up was finished, all five had turned out to be null results or regressions. But they only became possible and well-justified through this analysis.)

Related files

  • eval_action_dist — action counter plus transition matrix
  • plot_action_dist — plots
  • per-variant action-distribution data files
  • action-distribution and transition grid figures

Glossary

  • Stronger model — our earlier baseline with the multi-cell action; a notable step forward at the time.
  • Weaker model — a later attempt to improve on it that regressed.
  • action 7 (forward-until-collision) — the multi-cell “advance to obstacle” action, the main productive action.
  • Coverage contribution per action — how many new cells a given action covers, summed across all episodes.
  • Cells per call — how many new cells one action 7 call produces on average.
  • Transition matrix — how many times each “action X then action Y” sequence occurred.
  • Boustrophedon — a snake-shaped traversal pattern, a classical coverage approach.
  • Lawnmower — a classical (non-RL) lawnmower-shaped coverage heuristic.
  • Stochastic evaluation — the model picks actions by sampling from the probability distribution rather than always taking the most likely one.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR