What this is: we trained the model several times, each time with a different strength of a reward term that encourages smoother flight paths. We traced a curve of “reward strength vs coverage” and looked for the best setting. Why: an earlier run had hinted that one moderate strength roughly matched the baseline. Rather than guess, we decided to walk the whole curve first — a short sweep is far cheaper than committing to a long, expensive training run on the wrong setting.
Date: 10 May 2026 Machine: D2
1. What we wanted
Context: in earlier work we added a reward term that encourages smoother, less jagged flight paths. One control knob sets how strongly the agent is nudged away from zigzagging. Too weak, and the reward does nothing; too strong, and it overrides the actual coverage goal.
Hypothesis: there is a useful middle ground. Find it, and we know the right strength to carry into longer, more expensive experiments on this baseline.
Predictions:
- A peak somewhere in the middle, with results falling off at both extremes.
- Or a steady decline, meaning the weakest setting was already best.
- Or a peak toward the stronger end, meaning we should push harder.
- And, separately, that the effect might grow with strength on the hardest, narrow-corridor map.
2. What we did
We did several short training runs in a row, each tried a different strength of this reward term, sweeping from weaker to stronger. Each run trained for the same budget and took only a few minutes.
After each run we evaluated on five unseen maps, five episodes each, at three step limits (short, medium, long).
In parallel — while the training runs proceeded — we ran an action-frequency analysis on a previously trained model that used a weaker setting.
3. What we saw — the main picture
Coverage at the long limit vs the baseline
The curve had the classic shape: results improved as we increased strength, hit a clear peak at a moderate setting, and then collapsed sharply at the strongest setting. In short — a peak in the middle, drops at the edges.
The headline finding
At the best middle setting, long-limit coverage edged just above the baseline — a small but real gain. This was the first run in a while to clear the baseline by a noticeable margin on the long limit. Pushing the strength too far, by contrast, dragged coverage well below the baseline.
On narrow corridors — the effect strengthens, then breaks
On the hardest map with narrow corridors the benefit grew steadily up to the best middle setting, where the gain was several points over baseline at the medium limit — clearly the strongest corridor result of the sweep, and several times larger than the weakest setting produced. The drone passed corridors more cleanly and moved on to new areas faster. At the strongest setting the mechanism broke down and corridor coverage fell below baseline.
Episode length — not a simple trend
Episode length did not move in a single direction across the sweep. The weakest setting finished episodes noticeably faster — its main upside. Middle settings made the agent more cautious and slower, and the strongest setting never reached the completion threshold, exhausting the step limit every time — a qualitative breakdown.
4. What action-distribution analysis showed
We ran the weaker-setting model through an action-frequency analyzer and compared it with the baseline.
Headlines:
- The long “forward to obstacle” action did not grow meaningfully — this reward does not magically increase multi-cell motion the way a hard action constraint might.
- The model developed a strong preference for turning in a single direction. A plausible reason: consistent one-direction turning yields smoother heading changes and therefore less zigzag, which this reward favors.
- The “scan” action dropped off — it does not move the drone, so the smoothness reward gives it no signal.
Effectiveness of the long-forward action: at the short limit it covered fewer new cells per call than the baseline — the reward dislikes very long straight runs. At the long limit it covered more cells per call than baseline, because on long episodes the policy avoids re-crossing already-visited space. That is the mechanism behind the faster finishes.
5. What this means — decision
We used a simple decision rule for what to do next: if the best strength helped overall without hurting the hard corridor map, carry it forward into the next big experiment; if it helped overall but hurt corridors, combine it with an action constraint; if every strength above the weakest regressed, fall back to the weakest setting.
The best middle setting helped overall and gave the best corridor result, so we carry it forward as the baseline for the next experiment.
What carries forward:
- The best middle strength becomes the baseline for the next experiment.
- The single-direction turning preference is a potential weakness. Varying map orientation during training may push the model toward a genuinely smoothness-aware traversal rather than leaning on a lucky starting state — or it may break and regress.
- Per-map monitoring stays on, especially the hard corridor map and the map that this reward tends to hurt.
Cost-benefit: the short sweep gave us the whole curve and the optimal point. Without it we would have committed the next big run to the wrong setting and missed the small overall gain. A clear methodological win — plan the science, don’t guess.
6. What’s next
Next up is a longer experiment: take the chosen baseline and add data augmentation (rotations, horizontal and vertical flips, random crops) to the training maps on the fly. The aim is to reduce overfitting to a fixed set of training maps and to amplify the smoothness mechanism.
We also considered an input-importance test around neighboring strengths to fine-tune the peak, but the likely gain is small compared with the augmentation experiment, so it was deprioritized.
Glossary
- Smoothness reward — a reward term that encourages smoother, less jagged flight paths.
- Strength knob — the control that sets how strongly the smoothness reward nudges the agent away from zigzagging.
- Peak-in-the-middle curve — a “setting vs result” shape with a peak at a moderate value and drops at the extremes.
- Baseline — the current best model, trained without the smoothness reward.
- Corridor map — our “hard” map with narrow corridors.
- Single-direction turning preference — a strong, persistent bias toward turning one way, often an artifact of the starting state.
- Faster finishes — shorter mean episode length: the drone reaches the completion threshold and ends before the step limit.
- Sweep — a series of experiments varying a single parameter.
- Action distribution — how often the trained policy picks each action.
- Cells per call — the average number of new cells one long-forward action uncovers.
- Train-up / eval-flat — training metrics improve while evaluation stagnates, a sign of overfitting.
- Map augmentation — expanding the training set with geometric transformations of the maps.