What this is: The morning a brand-new training environment for indoor mapping came online, plus the first two “untrained pilot” runs we use as a yardstick. Why it’s here: Because one of those untrained pilots did far better than anyone expected, and that surprise quietly reshaped what success even means for the learning model.
Date: 2026-06-07 Ticket: rl-lab dev-log, entry 20a
Glossary
- Environment (the trainer): A small simulated world where the drone practices. Think of it as a flight simulator built for one specific lesson. Here the lesson is “explore an indoor space and draw a map of it.” The drone flies, its range sensors sweep the walls, and the unseen parts of the room gradually fill in.
- Baseline (the bar to beat): A simple, no-training solution we compare the learned model against. Like timing yourself against a stopwatch before hiring a coach. If the trained model can’t beat the bar, we fix the model rather than congratulate the bar.
- Random pilot: A “drone” that presses controls at random. The lowest, dumbest bar. It exists to show what pure luck achieves.
- Nearest-frontier (the crammer): A classic rule-based explorer from the late 1990s (Yamauchi, 1997). Its whole strategy is “always fly toward the closest edge of the unknown.” Reliable, but greedy about short distances, so it sometimes wanders inefficiently. This is the serious bar the model must beat.
- Multiroom layout: Generated “apartments” made of several rooms connected by doorways. Walls block the sensors, so you genuinely have to fly through a door to see what’s behind it.
- Steps-to-coverage: How many flight steps it takes to fill the map up to a target percentage. Our main measure of speed, not just completeness.
1. The new trainer is built
The big morning win: a fresh indoor-mapping environment is up and passing its full automated test suite, all checks green. The mental picture is “a room explored with a flashlight.” The drone flies around, its rangefinder beams trace the walls, and it earns credit for every fresh patch of the room it reveals for the first time.
This is the stage on which every later experiment will run, so getting it correct and well-tested matters more than getting it fast.
2. Two control pilots, before any learning
As planned, we ran two untrained pilots first, so we’d know exactly what we’re measuring against once training begins:
- The random pilot mashes controls with no plan.
- The nearest-frontier crammer uses that 1997 rule: always head for the closest edge of the unexplored area. This is the bar a trained model is obligated to beat.
3. The surprise (a useful one)
The original spec expected the random pilot to map less than 30% of a room. Reality came in much higher:
- About 70% coverage on a simple single room.
- About 40% coverage on a multi-room apartment with doors.
Why so high, and why this is not a bug: our rangefinder reaches roughly 6.4 meters, and a simple room is also about 6.4 meters across. The sensors read continuously, every instant of flight. So just by being in the room and spinning around randomly, the drone already sketches a large chunk of the map. The starting room is essentially mapped “for free.”
We deliberately did not cripple the sensors to make the number look more impressive. The simulated sensor matches the real TF-Luna hardware, and breaking that correspondence to win a prettier statistic would defeat the point of the simulator.
4. The task is still genuinely hard — here’s the real proof
The single-room number is misleading. The honest test is the apartment, where walls don’t see through and you must physically pass through doorways. We ran 50 attempts each and asked a simple question: did the pilot ever finish a near-complete map (up to 80%)?
| Pilot | Reached 80% coverage? |
|---|---|
| Random | 0 times out of 50 |
| Nearest-frontier (crammer) | 50 times out of 50, averaging about 450 steps |
The gap is the whole story. You do not stumble through doorways by luck; getting into the next room requires intent. Random flailing gets you the first room and nothing more. Because the original “less than 30%” target turned out to be unreachable in this setup, we flagged that the acceptance criterion should be revised to match what we actually observe.
5. What this changes about the training goal
The key realization: on a simple room, almost anything will reach a high coverage if you give it enough steps. So the contest was never really about how much of the map gets drawn. It’s about how fast.
That reframes the bar. The crammer reaches 80% of an apartment in roughly 450 steps. The model’s job is to do the same thing faster — by being smarter about where it flies. Instead of always heading to the nearest unexplored edge, a good model should head to the richest one: the doorway behind which an entire extra room is hiding, rather than a small dead-end nook nearby. That’s exactly where a learned policy can squeeze out an edge over the decades-old classic.
6. What comes next
Before committing serious GPU time to a full training program, we kick off a short trial run of an hour or two. It’s the cheap-probe-first discipline: prove the setup behaves on a small, inexpensive run before paying for the expensive one. If the trial looks healthy, the real training follows.
In short: the trainer works, the random pilot embarrassed our pessimistic assumptions, and the real benchmark turned out to be speed against a tough 1990s algorithm rather than raw coverage. A better starting bar, honestly measured, is a better foundation for everything that follows.