Imagine a small drone slipping into a building it has never seen — no map, no GPS, no overhead floor plan to consult. Smoke, darkness, dust: the kind of place where cameras give up. Its only sense of the world is a handful of distance beams reading how far the nearest walls are, and a thin steerable beam it can sweep around to probe ahead. The mission sounds almost trivial when you say it out loud: fly to a series of points inside the space, and don’t crash. Getting a learning agent to actually do that, honestly, from raw sensors and real-feeling physics, was the whole story of our A.1 sprint.
This is that story — the goal, the wrong turns, and the moment the drone went from a lab curiosity to a model we trust.
The goal: explore, don’t follow
It would have been easy to cheat. Hand the agent a perfect interior map and the task collapses into classical path-planning — a solved problem that needs no learning at all. We deliberately refused that shortcut. The point of A.1 was a drone that does not see the map: it flies on raw distance sensors, feels the inertia and tilt of its own body as it accelerates, and has to actively look around to understand a space it’s discovering in real time.
That single decision shaped everything that followed. We weren’t teaching a drone to trace a known route. We were teaching it to find its own way through the unknown.
The drone that crashed on purpose
The first scene, S1, is the tutorial level: a single straight corridor with a few waypoints to pass through in order. If the agent can’t learn this, nothing harder is worth attempting. We train using PPO, the workhorse algorithm of modern robotics RL, which nudges behavior toward whatever earns more reward in small, cautious steps.
And the drone failed S1 in the most maddening way imaginable. It wasn’t slowly improving and stalling out. It was crashing almost immediately — and it seemed content to do so. Episode after episode, it bolted straight into a wall within a handful of steps.
The instinct was to blame the corridor: maybe it felt too narrow, too claustrophobic. So we ran a clean test — same incentives, but in a wide-open hall with nothing to hit nearby. If geometry were the culprit, the drone should suddenly behave. It crashed just as fast. That was the tell. The problem wasn’t the world. The problem was how we were scoring the drone.
This is one of the oldest traps in reinforcement learning, and it’s worth understanding because it’s so counterintuitive. We had been gently discouraging the agent on every single step — small nudges toward smoother, calmer, more efficient flight. Early in training, when its movements are noisy and exploratory, those tiny discouragements pile up faster than it can earn anything good. Every moment of life felt slightly costly. Meanwhile, crashing ended things with a single, finite cost. Do the grim arithmetic from the drone’s point of view: if living drains points forever and dying costs you only once, then dying sooner is cheaper. The agent had rationally concluded that the best way to stop losing was to stop existing. (We wrote up this failure mode in depth in when the drone games the metric.)
Stripping it back to essentials
The cure was subtraction, not addition. We tore the incentive down to the bare minimum: make progress toward the next waypoint, complete the mission, and don’t crash. Every secondary “fly nicely” signal got switched off.
With that, the drone’s preferences flipped into the right shape — navigating beat hovering, and hovering beat crashing fast. Suddenly, flying the route was genuinely the most rewarding thing it could do, and it climbed straight toward the goal. On S1 it hit 100% of waypoints with zero collisions, flying the corridor a touch tighter than a hand-built classical planner. A small touch helped too: we started each episode with the drone roughly facing its first waypoint, rather than spawning it nose-first into a wall.
The lesson became a guiding principle for the rest of the sprint: teach navigation first on the simplest possible reward, then layer the polish back in gradually — one refinement at a time, only after basic flight is solid. Pile every demand on at once and you smother the very exploration the agent needs to learn to fly at all. (See dev-log 48 for the full diagnosis.)
Cracking the apartment
Four of the easier scenes fell quickly to the minimal recipe; one larger hall needed a longer training run. Then we hit the holdout: the apartment — several rooms strung together by tight doorways. The drone failed it in two opposite ways. Trained aggressively, it had the energy to reach the doors but no instinct to slow down, and clipped the frames. Trained to be cautious from scratch, it never learned to thread a doorway at all — it just froze and timed out.
Two ideas, combined, finally cracked it. The first was a warm-start: instead of training the apartment drone from a blank slate, we started it from a policy that had already mastered a simpler two-room scene. Like teaching someone to drive on quiet streets before handing them the keys in city traffic — they arrive already knowing the basics. The second was a near-wall caution: when the drone is racing straight at a wall, ease off — unless there’s a doorway there it’s meant to fly through. Elegantly, that exception falls out naturally, because the path through a door is open space, so the caution simply doesn’t fire there.
Neither ingredient worked alone. Warm-start gave competence but no restraint; the safety instinct from scratch gave restraint but nothing to build on. Only together — with the caution tuned to just the right firmness, neither timid nor reckless — did they click into place: a clean 100% of waypoints, zero collisions, flying smoothly through every doorway. With the apartment solved, we had a known recipe for the whole A.1 scene set. (Dev-log 51 has the full ladder of attempts.)
Learning to actually scan
There was still something hollow about it. Watching replays, we noticed the drone was largely ignoring its steerable scanning beam — the close-range distance sensors alone were enough to stumble through. But a drone meant to map unknown spaces can’t fly half-blind. So we made scanning matter: we shortened the reactive sensors to a close emergency bumper, so the only way to know what lay further ahead was to sweep the beam and look. Flying into an un-scanned patch of space now carried a cost.
The drone learned, in effect, to look before it leaps. It began actively sweeping its beam across the path ahead, building an understanding of each room as it moved — and the rate at which it charged into un-scanned space dropped roughly fourfold. It still cleared all the scenes, doorways included, but now it was genuinely perceiving its way through rather than getting lucky. Active scanning had become a learned skill, not an afterthought. (More in active scanning as a learned skill.)
Is it real, or just luck?
One flawless model proves a recipe can produce a great drone. It does not prove the recipe reliably does — the random dice at the start of training might simply have landed well. Before promoting this model to production, we asked the obvious question: does it repeat?
So we re-ran the identical training recipe from scratch with three more random seeds, then put every resulting model through every scene — a full grid of independent exams. The result: every model, on every scene, completed 100% of tasks with zero collisions. A curious wrinkle — the models flew nearly identical routes, which briefly looked suspicious — turned out to be honest: they shared a common foundation and the optimal flight path was simply unique, while each model still differed in how it swept its scanner around. The fingerprints of the model files confirmed they were genuinely distinct. The success wasn’t a fluke. The recipe reproduces. (Multi-seed validation and dev-log 53 tell the full story.)
With reproducibility settled, the scan-v1 model was promoted to the production model of sprint A.1.
What it means, and what’s next
Step back and the arc is clean. We started with a drone that crashed on purpose because we’d accidentally taught it to. We fixed the incentive down to its honest essentials and watched it learn to fly. We gave it a warm-start and a sense of caution and it learned to thread doorways. We forced it to look, and it learned to scan. And finally we proved, across many independent runs, that none of it was luck.
That last step is the quiet milestone. The difference between “one promising result” and “a recipe we trust” is the difference between a demo and a foundation — and A.1 is now a foundation. From here, the work shifts from new capability to finishing touches: making the motion smoother and gentler, and moving a step closer to real hardware in higher-fidelity simulation. But the hard question — can a drone learn to navigate and scan a space it has never seen, honestly, from raw sensors? — has an answer now. Yes. And we can build the next sprint on top of it.