claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

05 — Classical lawnmower heuristic: when a simple snake beats RL

We built a classical coverage algorithm with no learning - a simple snake that bypasses obstacles via BFS - to serve as an honest reference point for RL.

stablerl-labupdated 2026-05-11T00:00:00.000ZClaudeDroneRLDevLog

What this is: a classical coverage algorithm (no learning, no networks) — a simple snake with BFS-based obstacle bypass. We needed it as a reference for an honest comparison with RL. Why: several runs in a row produced roughly 17–19% coverage at 1000 steps. Is that a lot or a little? Without a classical reference there was no way to tell. We ran the heuristic and got 26.1% — RL came out about 1.4× behind the simple approach.

Date: 8 May 2026 Machine: D2 (but the heuristic runs on CPU — no training needed)


1. What we wanted

In a research task, without a baseline you cannot say whether “19%” is a lot or a little.

Lawnmower is a standard classical coverage algorithm:

  1. Walk along a row until hitting a wall.
  2. Step down one row.
  3. Walk back the other way until hitting a wall.
  4. Step down again. Repeat.
  5. When obstacles appear, use BFS to reach the nearest next unvisited point on the route.

With known geometry this gives the best achievable result, so any RL policy should at least be able to match it.

2. What we did

We implemented a compact lawnmower evaluator (around fifty lines):

  • A sweep across rows with direction reversal at each wall.
  • BFS bypass around obstacles to the nearest next free cell.
  • It stops once no reachable unexplored cells remain.

We ran it on the same five maps, five start positions each, with step budgets of 1000, 3000, 5000 and 10000.

3. What we got

Step limit Lawnmower RL (best run) Gap
1000 26.1% 19.0% −7.1 pp
3000 68.9% 41.5% (extended eval) −27.4 pp
5000 98.7% 49.6% −49.1 pp
10000 100% 65–73% −27 pp

RL trailed the heuristic at every step budget — about 1.4× behind at 1000 steps and 1.7× behind at 5000.

Across maps the lawnmower was very consistent: 25.6–26.9% at 1000 steps on all five maps, while RL showed a spread of roughly ten percentage points.

4. What this means

Reframing the mission.

The goal of “coverage ≥ 85%” now reads in two ways:

  1. On an unknown map (our task): at minimum, catch up to the lawnmower’s 26% at 1000 steps.
  2. On long episodes (5000+ steps): approach the 98.7% mark.

Why RL trailed:

  • The lawnmower effectively knows the map, since the sweep follows its coordinates. RL only sees local distances from its rangefinders plus a visited grid, so it has to rediscover the structure on its own.
  • The lawnmower follows an explicit plan, whereas RL takes policy steps without long-horizon planning.
  • A stochastic policy combined with a collision penalty means RL spends some steps unproductively.

Even so, RL is the right tool when:

  • The map is not known ahead of time (our setup with domain randomization fits this).
  • Generalization matters — RL learns, whereas the classical method is hand-coded.
  • The scene has dynamic elements (target detection, moving obstacles), where a learned policy outperforms a fixed sweep.

The headline takeaway: the lawnmower is the reference we need to reach. Until this point we had largely been measuring ourselves against ourselves, overlooking how far ahead a simple heuristic actually was.

5. What’s next

  • The lawnmower result is now recorded as a reference in the training table.
  • A longer-horizon retrain is running in the background so we can watch the gap.
  • A multi-cell forward action is rising in priority. The lawnmower is effective because it makes long, uninterrupted passes; if RL can move forward until a collision, it may close more of the distance.

Related files

  • The lawnmower evaluator implementation.
  • The metrics file with the measured numbers.
  • The training table, with the lawnmower row added.

Glossary

  • Lawnmower — a coverage algorithm of the sweep-and-turn family. Simple and effective.
  • Boustrophedon — the formal name for the snake-shaped traversal.
  • BFS (Breadth-First Search) — used here to bypass an obstacle.
  • A* — a more advanced path-finding algorithm that uses a heuristic.
  • Upper bound — the best achievable result; the lawnmower sets it on a known map.
  • Coverage — the fraction of free cells visited.
  • DR (Domain Randomization) — training across many random maps so the model generalizes.
  • Wake-up call — the moment you realize how badly you had misread progress.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR