What this is: we added a rule so that when the drone is in a good position — with enough clear cells ahead — only a small subset of actions is permitted: turn left, turn right, or the productive forward move. The drone is free to choose any of them. Why: earlier analysis showed that the forward-until-collision move is responsible for almost all of the coverage, yet the policy only used it a small fraction of the time. We hoped that nudging the policy toward it would yield more coverage.
Date: 8 May 2026 Machine: D2
1. What we wanted
Earlier analysis suggested that the forward-until-collision move — a multi-cell motion that sweeps ahead until it hits an obstacle — accounts for nearly all of the coverage the agent achieves, even though the policy rarely selected it. The thought was that if the agent used this move more often, coverage might improve without any architectural changes.
The idea was to hint the policy with a mask: whenever there is a clear opportunity to move forward, restrict the choice to turning left, turning right, or the forward move, and block everything else. The reasoning was that turns are safe maneuvering options while the forward move is the productive one.
We expected the forward move to be selected much more frequently, and coverage to rise at both the short and long step horizons, narrowing the gap to the classical baseline.
2. What we did
We built a masked variant of the environment that applies the restriction whenever enough cells are clear ahead. For the algorithm we used the action-masking variant of PPO from the sb3-contrib library, which supports blocking actions on a per-step basis.
Training ran for one million steps in roughly 7.4 minutes. The run looked clean, with a high explained-variance score and a healthy final reward.
3. What we saw
Coverage — essentially unchanged, with a slight regression
| Step limit | Mask | Baseline EXP-7 | Δ |
|---|---|---|---|
| 1000 | 56.93% | 56.10% | +0.83 (noise) |
| 3000 | 84.80% | 86.18% | −1.38 |
| 5000 | 92.21% | 93.75% | −1.54 ⚠ |
Action distribution — the main observation
| Action | Mask | Baseline | Δ |
|---|---|---|---|
| turn left | 47.1% | 23.0% | +105% ← explosion |
| turn right | 10.3% | 35.6% | −71% |
| forward move | 10.2% | 10.9% | −6% (did not grow) |
| others | 32.5% | 30.5% | ≈ |
The headline: the forward move did not become more frequent. Under the mask, the drone overwhelmingly chose to turn rather than to use the productive forward action.
Turning in one direction roughly doubled in the process. So out of the permitted set, the drone simply preferred to turn. The shift between the two turn directions is best read as an initialization quirk rather than a meaningful signal.
Why the numbers came out this way
When only a small set of actions is permitted and exploration is modestly encouraged, the choice among them tends toward a roughly even split. The mask itself was only active perhaps a third of the time — whenever the path ahead happened to be clear enough. Combining an even split among the permitted actions with how often the mask was active gives an expected frequency for the forward move of only a little over ten percent, which lines up almost exactly with what we observed.
In short, the predictions did not hold up.
4. What this means
The main lesson is that permitting an action is not the same as requiring it. A soft mask permits the productive forward move, but it does not force the agent to take it. The agent optimizes its return, not the frequency of any particular action.
There is a reward asymmetry at play. Turning is a safe move that carries no penalty, while the forward move carries collision risk and is penalized if it crashes. Without an additional incentive attached to the forward move, the safe strategy of turning simply wins.
This is consistent with the literature on action masking. Masking is meant for cases where an action is genuinely invalid under the rules of the task. When every action is technically valid and the mask only encodes a preference, it distorts the policy rather than guiding it. Community discussion echoes the same point: this style of masked PPO works best for ruling out illegal moves, not for coaxing the agent toward a favored option.
This points clearly toward what to try next. Since a soft restriction did not work, a stricter restriction is the natural follow-up — permitting only the forward move in a good position. (As a preview: that approach also produced a null result, covered in the next entry. The agent learned to avoid good positions altogether so it would never fall under the mask, a textbook case of optimizing the proxy instead of the goal.)
5. What’s next
We see several directions worth exploring next. A stricter mask permitting only the forward move is the immediate follow-up, though it risks getting the agent stuck in awkward geometry. A recurrent policy is another option and is less prone to being gamed. A hierarchical approach is also appealing, where a higher-level option such as “sweep until collision” provides structural pressure without relying on a mask at all.
Related files
envs/drone_2d_env_masked.pytraining/configs/ppo_masked_multicell.yamltraining/{train,eval}_masked.py
Glossary
- MaskablePPO — a variant of PPO that can block actions. From the
sb3-contriblibrary. - Soft mask — a soft restriction that allows a small subset of the available actions.
- Hard mask — a strict restriction that allows only a single action.
- action_masks() — a function in the environment returning, for each action, whether it is currently allowed.
- ActionMasker — a wrapper from
sb3-contribthat hooks the environment so masks reach the algorithm. - cells_free_ahead — an environment-side helper that counts how many cells are clear along the current heading.
- forward move (forward-until-collision) — a multi-cell action that advances forward until it meets an obstacle.
- Lawnmower — the classical, non-RL coverage heuristic used as our reference; it reaches about 98.7% at 5000 steps.