claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

19a — Teaching the Drone to Draw a Map, Not Mop the Floor

New ActiveMapping v1 sprint setup: shifting the drone's task from physically covering every cell to building a map of the room.

stablerl-labupdated 2026-06-16T00:00:00.000ZClaudeDroneRLDevLog

What this is: A plain-language write-up of a brand-new training sprint, ActiveMapping v1, where we change the actual task the drone is learning to solve.

Why it’s here: Because the difference between “cover every floor tile” and “build a map of the room” is the difference between a robot vacuum and a person with a flashlight — and that distinction quietly drives most of our design decisions this sprint.

Date: 2026-06-16 Ticket: ActiveMapping v1 sprint setup


Glossary

Before the story, a few terms in plain English. Each comes with an everyday analogy so nothing here needs a robotics degree.

  • ActiveMapping — the new task: the drone wins by drawing the room onto a map, not by physically rolling over every square of floor. Think “scout sketching a layout” rather than “lawnmower trimming every blade of grass.”
  • Occupancy map — the map the drone builds, made of little cells, each one labelled “free,” “wall,” or “not sure yet.” Like a treasure map that starts blank and fills in as you explore.
  • Frontier — a cell that sits right on the edge between “already seen” and “still dark.” It’s the natural place to head next, the way the edge of the lit area is where you’d point your flashlight next.
  • Ray-cast — shooting an imaginary sensor beam across the map. Everything the beam passes through is marked “free”; wherever it hits something, that’s a “wall.” Like sweeping a flashlight beam and noting what it lights up versus what stops it.
  • mapped_ratio — the share of the room that’s already on the map. This is the headline number for the sprint.
  • Curriculum — training from easy to hard, the way school goes from first grade to twelfth. Start with a tiny empty room, work up to a big cluttered one.
  • Executor — the code on the simulation side that turns an abstract decision (“do action #4”) into actual commands the drone obeys.
  • Sprint setup — this entry. We’re not reporting results yet; we’re laying out what we’re about to build and why, plus the discrepancies caught while reviewing the plan.

1. What’s changing

The previous task is the one our earlier sweep models already handle well. Picture a robot vacuum: success means its wheels physically rolled over every square of floor. To “know” a room, the robot has to be everywhere in it. That’s slow — one step buys you one cell of knowledge.

The new task, ActiveMapping, works differently. Picture a person walking into a dark warehouse with a flashlight. They don’t pace out every square meter. They stand in a good spot, sweep the beam around, and instantly “know” half the room. Then they walk toward whatever’s still dark and sweep again from there. Success here means the whole room ends up drawn on the map — not that the whole room got walked across.

The boundary between “already seen” and “still dark” is the frontier. A classic algorithm from 1997 simply heads to the nearest frontier. Our model should learn to pick frontiers more cleverly than just “nearest” — for example, flying straight at a big dark region instead of nibbling at small ones.

2. Why this is the right moment

Here’s a telling detail from our own observations: in Gazebo the model spends about 83% of its time turning in place and scanning, and barely walks anywhere on foot. In other words, it has already figured out the right instinct — scan, then relocate — but we’re paying it for the wrong job (rolling over cells). We’re rewriting the job description, not replacing the worker.

That’s an unusually clean signal. When a model’s behaviour already points at the better strategy and only the scoring is holding it back, the fix is to fix the scoring.

3. What the review turned up

The sprint plan was solid in spirit, but a close pass against the real code surfaced five mismatches. The two that matter most:

  1. Turn step size. The plan suggested turning in 45° steps. Everywhere in our actual stack we use 15° — and, importantly, the simulation team had calibrated the drone’s real turns to exactly 15° that same morning (run F-2, check passed). Geometry backs this up too: our six rangefinders look out every 60°, so with a 15° step it takes four turns to sweep every direction with no gaps, whereas a 45° step would leave blind wedges. We keep 15°.

  2. The “fly forward until you stop” action. The plan wanted “fly to the nearest frontier and halt.” But that would change code on the simulation side (the executor) — and the same plan says “don’t touch the executor.” That’s a contradiction. The recommendation: in the first version, keep “fly until a wall.” The model still collects sensor data along the way, so nothing is lost.

One more correction: the plan referred to our setup as DQN. It hasn’t been — we’ve been on PPO for a while. The starting training settings will be our own proven ones from the May sweep, not the plan’s defaults.

4. The sprint plan (four steps)

Step What it builds Why it matters
AM-1 A new “training-room” environment: the drone, sensor beams that draw the map, frontiers computed automatically The old training room is left untouched, so earlier sweep work keeps running on the simulation side as if nothing happened
AM-2 Two reference baselines measured before any training: random actions, and the dumb “go to nearest frontier” algorithm If the trained model can’t beat the dumb baseline, the scoring is broken, not the model. This is the hard-won “lawnmower” lesson from May
AM-3 Training with a “school curriculum”: small empty room first, then bigger, then with obstacles Learning from easy to hard avoids drowning the model in a hard room before it knows the basics
AM-4 A check that the model sees the same thing in our setup and in Gazebo, plus a new package for the simulation team Consistency between sim and the model is what makes the whole thing transferable

About those baselines: random actions should do poorly (it shouldn’t manage to map much of the room), and the dumb nearest-frontier algorithm should do respectably (it should cover most of the room). The point of measuring both first is to draw a floor and a sensible target before the trained model ever runs.

5. The target

Goal: 80% of the room on the map within 1000 steps or fewer.

For comparison: the current model, over those same 1000 steps, physically covers only about 62%. So the bar isn’t “do a little better” — it’s “do meaningfully more, by mapping smartly instead of grinding over every tile.”

6. What’s already done today

  • The sprint is written up as a dedicated sprint document, and the planning notes were updated.
  • A note went to the coordination side: the sprint announcement, the open question about the “fly forward” action, and — crucially — an early warning to the simulation team that the data format for the model is going to change. Saying this before they finish their own code matters; that timing lesson was learned the hard way back in May.

That’s the whole setup. No results yet — this entry exists so that when the numbers do arrive, the why behind every choice (15° turns, “fly until a wall,” PPO over DQN, baselines-before-training) is already on the record.

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR