What this is: we added a reward that gently discourages the drone from wiggling back and forth, so its path through a space stays smoother. Why: several earlier changes in a row had all hurt our results. This was the first change that did not make things worse — not yet a clear win, but the underlying idea clearly does something useful.
Date: 9 May 2026 Machine: D2 (RTX 5070)
1. What we wanted to check
We were working on the same kind of task that appears in recent coverage-planning research: a drone exploring and covering an indoor area without a prior map. One recurring idea in that literature is to gently discourage jagged, back-and-forth movement so the covered region grows in a cleaner, more compact way. The reported gains there were large, so we wanted to know whether the same idea would help on our setup.
The intuition is simple: instead of only rewarding new ground covered, we also apply a mild discouragement whenever the drone’s recent movement makes the boundary of the covered area more ragged. The drone is nudged toward steady, purposeful sweeps rather than nervous wiggling.
There is a single knob that controls how strongly we discourage that wiggling. A light setting suits open exploration of unknown space (our case); a heavy setting suits tidy, methodical sweeping of an already-known region. We chose the light setting for this run.
Going in, we sketched a range of possible outcomes — anything from a meaningful coverage gain, through no real change, to a small regression if the discouragement turned out to be too aggressive for our environment.
2. What we did
Implemented — a short, focused change:
- A new training environment that extends our main environment and adds a per-step smoothness signal.
- The smoothness signal looks at how ragged the boundary of the covered region is before and after each step, and turns any increase in raggedness into a small negative contribution to the reward.
- The training configuration was kept identical to our previous best run, except for the few new settings that control the smoothness signal.
Training — roughly eight minutes on one million steps:
- Throughput dropped by about a quarter, because the smoothness signal has to be computed on every step.
- Average reward per episode dipped by a few percent. That is expected, not a defect — we are deliberately spending a little reward in exchange for smoother trajectories.
- Training was stable: no numerical blow-ups, no crashes.
Evaluation — a few minutes:
- Five maps the model had never seen during training, five runs each, at three different step budgets (1000, 3000, and 5000 steps).
3. What we saw
Average coverage — essentially unchanged
| Step limit | Smoothness run | Best prior run | Δ |
|---|---|---|---|
| 1000 | 54.92% | 56.10% | −1.18 |
| 3000 | 85.72% | 86.18% | −0.46 |
| 5000 | 93.88% | 93.75% | +0.13 |
At the largest step budget, coverage came out the same as before, but the drone reached it using about 123 fewer steps on average. Same outcome, achieved faster.
Looking map by map is where it gets interesting
The maps differ a lot — open spaces, narrow corridors, semi-enclosed rooms. Averaging across all of them hides an important pattern.
On maps where near-complete coverage is reachable, the drone tended to finish sooner:
- On one open map it saved several hundred steps, around a tenth of its time.
- On another it saved a smaller but still clear chunk of time.
On maps where near-complete coverage is not reachable, the drone tended to cover slightly more:
- On the narrow-corridor map, coverage edged up by about a point at the larger budgets.
- On the most open map, coverage was flat at the largest budget, but it dropped noticeably at the middle budget — the one clear regression we observed.
A prediction worth recording
Before the run, Aleks predicted that the corridor map would benefit from discouraging wiggling. The opposite seemed more obvious to me: a corridor already forces a roughly linear path, so there should be little wiggling to discourage and therefore little to gain.
In practice the corridor map did benefit, and the reason became clear afterwards. In a tight corridor the drone tends to loop and double back against the wall, hoping to pick up cells it might have missed — and that doubling back is exactly the wiggling we discourage. With the smoothness signal in place, the drone moves through the corridor more cleanly and pushes on to new ground sooner.
The lesson I took away: I had been reasoning about the final shape of the covered area, a static picture, when I should have been reasoning about how that shape changes step by step, a dynamic one. I noted that down for next time.
4. What this means
The mechanism works, but it shows up in two different ways:
- On easier maps, it saves time while reaching the same result.
- On harder maps, it adds a little coverage, on the order of a point or two.
Because these two effects pull in different directions, the overall average ends up looking like no change at all — which is why the headline comparison reads as flat.
This was the first time in a run of experiments that adding something did not make results worse, and it happened by directly reusing a known idea from the literature rather than by hand-tuning settings.
The cost was modest: slower training and a small dip in raw reward, both of which we knew about in advance.
It is also worth being honest about why we did not see the dramatic gains reported elsewhere. Our prior run was already a strong starting point, leaving far less headroom for a big jump, and we deliberately used the light setting of the smoothness knob rather than the heavy one that produced the largest published gains in a different, tidy-sweeping setting.
5. What comes next
A couple of directions are open for discussion:
- Trying a stronger setting of the smoothness knob with the same input-importance test, to see whether it amplifies the corridor gain — at the risk of hurting the most open map further.
- Folding this together with an upcoming experiment that increases map variety through rotations and reflections, which may combine well with smoother trajectories.
For now the smoothness reward stays in the toolkit as a base ingredient for future combinations.
Related files
- The training environment that applies the smoothness penalty.
- The training configuration for this run.
- The stored results for this experiment.
- Our internal notes on the coverage-planning literature this idea came from.
Glossary (terms from this dev-log)
- PPO — Proximal Policy Optimization, our main reinforcement-learning algorithm.
- Trajectory smoothness — informally, how clean and steady the drone’s path is, versus jagged back-and-forth movement.
- Percentage point (pp) — a difference in percentages. Going from 86% to 87% is one percentage point.
- Parity — a difference small enough that we treat the two results as effectively the same.
- Baseline — our reference point, the best prior result.
- Eval maps — the five maps used to test the model, none of which appeared in training.
- Stochastic eval — evaluation where the model samples actions by probability rather than always taking the most likely one.
- Input-importance test — a controlled run where we change one factor on its own to understand how much it matters.