What this is: we added a rule — whenever a clear stretch of open space appears in front of the drone, it is required to move straight forward and every other control is locked out. Why: in the previous run the drone was allowed to either go forward or turn, and it preferred turning. So this time we tried to force the issue. The result: the drone learned to steer clear of the very positions that would trigger the rule, and coverage collapsed by half.
Date: 8 May 2026 Machine: D2
1. What we wanted
Context: an earlier analysis showed that the multi-cell “move forward until you hit something” action delivers almost all of the drone’s coverage, yet it was chosen only a small fraction of the time. The intuition seemed obvious: encourage the drone to use it more often and coverage should rise.
Previous run (soft approach): in an open position the drone was allowed to choose among moving forward or turning either way. It still preferred the turns, and use of the forward action did not increase. No improvement.
This run (forcing approach): in an open position the drone could only move forward; every other control was unavailable.
An “open position” meant a clear run of several cells directly ahead.
We expected the forward action to be used far more often, coverage at both short and long step budgets to climb noticeably and close the gap to the classical baseline, and we anticipated a risk that the constraint might leave the drone stuck in awkward geometry.
2. What we did
Implementation: we extended the environment, training, and evaluation scripts with a switch to select between the older soft approach and the new forced approach.
Quick check (short run): the early-training reward looked much higher under the forced approach than under the soft one, which seemed encouraging — as if forcing forward motion paid off immediately.
This turned out to be a trap. The early reward signal was misleading.
Full training ran for the usual budget in roughly eight minutes. The final score came out far lower than both the soft approach and the baseline — down by about half relative to baseline. That was the first sign something was badly wrong.
3. What we saw
Coverage collapsed
| Step limit | Forced | Soft | Baseline | Lawnmower |
|---|---|---|---|---|
| 1000 | 29.90% | 56.93% | 56.10% | 26.1% |
| 3000 | 34.15% | 84.80% | 86.18% | 68.9% |
| 5000 | 36.73% | 92.21% | 93.75% | 98.7% |
At 5000 steps coverage was only 36.7% — the drone gained barely a few points for four times the time budget. That is not coverage, it is wandering.
Across all five maps the regression was uniform, roughly −55 to −61 points. Not a single map worked.
What the action breakdown showed — the key finding
| Action | Forced | Soft | Baseline | Δ vs baseline |
|---|---|---|---|---|
| forward-until-collision | 13.0% | 13.6% | 12.7% | 0% (!!!) |
| rotate one way | 18.8% | 43.0% | 32.0% | −41% |
| rotate the other | 20.7% | 27.3% | 34.3% | −40% |
| scan | 11.9% | 2.7% | 4.1% | +190% |
| strafe right | 11.9% | 3.6% | 5.6% | +113% |
| 1-cell back | 10.5% | 4.1% | 5.3% | +98% |
| 1-cell forward | 4.5% | 0.9% | 1.0% | +350% |
The telling point: use of the forward action did not grow at all — the same fraction as without any rule.
What did grow: scanning, strafing, and tiny one-cell movements. What do those share? None of them brings the drone into an open position with a clear run of cells ahead.
The drone learned to avoid open positions
This is Goodhart’s law in action:
“When a metric becomes a target, it ceases to be a good metric.”
We rewarded one thing — how often the forward action fired. The drone optimized something else entirely: spending as little time as possible in the positions that would force it. The two are not in conflict, because once it was in an awkward position it was free to do anything.
Even when the forward action fired, it covered half as much ground
| Forced | Baseline | |
|---|---|---|
| Number of calls | 3247 | 3182 |
| Mean cells per call | 8.10 | 16.14 |
| Median | 1.0 | 12.0 |
The forced approach triggered the forward action the instant a short clear stretch appeared, only to meet a wall right after, so no long sweep ever developed. The baseline, with no rule at all, chose the forward action in much better spots — deep inside corridors — covering on average about twice as much ground per call.
4. What this means
The optimistic predictions were all wrong. We did, however, confirm the risk we had flagged, though in an unexpected shape: the drone did not get stuck — it deliberately steered around the conditions that would trigger the rule.
Methodology lesson: when an intervention restricts which actions are available, always check whether the agent can simply avoid the situations that bring the restriction into play.
How to spot this kind of gaming:
- If the targeted behaviour does not increase but coverage falls, the agent is probably gaming the trigger condition.
- The symptom is a rise in “safe” actions — scanning, strafing, micro-movements — that never lead into the triggering situation.
The literature agrees:
- “Specification gaming in RL” (arXiv 2310.02842) catalogues cases where an agent finds a loophole. Ours fits the pattern exactly.
- “Invalid Action Masking” (arXiv 2006.14171), section 4, argues that masking should apply only to genuinely impossible actions; expressing a mere preference distorts the policy. In our case the forward action in an open position was perfectly valid — we simply forced it — which is exactly the kind of distortion that paper warns against.
This whole line of work is now closed:
- The soft approach produced no improvement.
- The forced approach was a catastrophe.
- We also dropped the idea of simply rewarding the forward action with a bonus, because it carries the same fundamental gaming risk.
5. What’s next
We are pivoting to directions that do not impose an explicit constraint the agent can dodge:
- A recurrent (memory-based) policy — our top candidate, since it adds no explicit constraint to game.
- A hierarchical approach where a higher-level option means “sweep until collision a number of times” — structural encouragement without a lock-out.
- Adding distance-to-frontier information to the observations — extra context rather than a constraint, and risk-free.
Related files
- environment module with the soft/forced switch
- the corresponding training configuration
- the training, evaluation, and action-distribution scripts
- the experiment folder for this run
Glossary
- Action masking — a technique for blocking forbidden actions. Blocking genuinely impossible actions is fine; using it to nudge the agent toward a preferred choice is risky, because the agent tends to game it.
- Forcing approach — a strict restriction: in an open position only one action is available.
- Soft approach — a looser version that permits several actions and lets the agent choose.
- Goodhart’s law — from an economist writing in 1975: “When a metric becomes a target, it ceases to be a good metric.” Frequently cited in RL.
- Specification gaming — the agent optimizes exactly what you specified rather than what you actually wanted. A standard problem in RL.
- Trigger condition — the situation that activates a rule (here, a clear run of several cells ahead).
- forward-until-collision — a multi-cell action that moves forward until an obstacle stops it. The main driver of coverage in earlier runs.
- PPO — the reinforcement-learning algorithm used here; we relied on a variant with action-masking support.