Learn to drive in an empty lot first
Nobody learns to drive in rush-hour traffic. You start in an empty parking lot, where the only thing to hit is the curb. You practice steering, braking, and judging distance with zero pressure. Once those become second nature, you graduate to quiet streets, then busier ones, and only much later do you merge onto a packed highway. Each stage builds on the skills from the one before it.
This is exactly the idea behind curriculum learning for an autonomous drone. Instead of dropping a freshly initialized agent into the hardest possible environment and hoping it survives, you teach it the way you’d teach a person: easy lessons first, hard ones later. The skills learned on simple scenes — hold a heading, keep distance from a wall, reach a target — transfer upward and make the difficult scenes learnable. The name comes straight from a school curriculum: you don’t start a first-grader on calculus.
For a flying robot navigating indoors, getting that order right is the difference between a policy that learns smoothly and one that thrashes around, never building a stable foundation.
Why easy-to-hard actually helps
A reinforcement-learning drone learns by trial and error. It flies, it succeeds or fails, and the algorithm — in our case PPO (Proximal Policy Optimization) — nudges its behavior toward whatever worked, cautiously, a little at a time. The catch is that early in training the drone is essentially random. In a brutal environment, almost every attempt ends in a crash, and the agent gets almost no useful signal about what good flight looks like. It’s drinking from a fire hose of failure.
Start it in a simple scene instead — a single straight corridor — and success becomes reachable by accident often enough that the agent can learn what success feels like. Once it reliably handles a corridor, an L-shaped turn is a small step up. A scene with one doorway is a modest stretch beyond that. By the time it faces a genuinely hard layout, it already owns the primitives it needs. Each rung of the ladder is close enough to the last that the drone can climb it.
Open rooms are easy. Apartments are not.
To make this concrete, we mapped a difficulty landscape across seven indoor scenes — a mental picture of how hard each one is relative to the others. Picture a terrain map: flat plains mean “easy to learn,” and steep mountains mean “the agent keeps failing here.” Most of our scenes turned out to be gentle plains. One was a stubborn mountain.
The plains were the single-path scenes: a straight corridor, an L-shaped corridor, a zig-zag, and even a large open hall with columns. The corridors were solved cleanly and quickly — there’s essentially one way to go, so “head toward the goal” is almost always the right move. The big open hall took more flying time, simply because there was more area to cover, but it got there. Open space is forgiving: there’s room to recover, few ways to get trapped, and a wall is something you drift toward, not something that suddenly boxes you in.
The mountain was the apartment — a multi-room layout connected by narrow internal doorways. This is where indoor autonomy gets genuinely hard, and the reasons are worth naming:
- Doorways are choke points. An opening roughly 1.4 m wide is a tiny target. A naive “get closer to the goal” instinct pulls the drone straight at the wall beside the door instead of threading the gap.
- Dead ends and recovery. In a multi-room layout, the shortest line to the target often runs through a wall. The drone has to back out, reroute, and recover — skills an open room never demands.
- Compounding decisions. Three or four rooms mean three or four doorways in sequence. One clean pass isn’t enough; the drone has to repeat the hard maneuver again and again.
The most instructive result was counter-intuitive. Early in apartment training, the policy was cautious: it timed out without reaching the far rooms, but it never crashed — safe and useless. With far more training, it grew aggressive, chasing distant targets and clipping the doorframes — collisions climbed sharply. More raw experience made it worse in the way that matters. Picture a delivery driver who, with more practice, decides to take corners faster to save time and starts scraping the gateposts. That’s the apartment in a nutshell: the doorway is where speed and safety collide, and brute-force training alone won’t resolve it.
How to build a curriculum: a checklist
If you’re sequencing scenes for an indoor navigation agent, here’s the practical shape that worked for us:
- Rank scenes by structure, not by gut feel. Single-path layouts (corridors) are easiest; open spaces are next; multi-room layouts with doorways are hardest. Let geometry, not intuition, set the order.
- Keep the objective identical across scenes. Don’t quietly help the agent on the hard ones. If the goal stays the same, any difference in outcome is about the scene, which is exactly what you want to measure.
- Evaluate the same way everywhere. Use a fixed, deterministic set of evaluation runs per scene so the numbers are apples-to-apples and you can actually see the difficulty landscape.
- Warm-start the hard scenes from related solved ones. A policy that already handles one doorway is a far better starting point for a many-doorway apartment than a cold start.
- Encourage clean motion, not just goal-chasing. On the easy scenes, reward smooth flight so the drone learns to pass through a gap rather than lunge at a goal. That habit pays off at every doorway.
- Treat “train longer” as a hypothesis, not a cure. If more experience makes a scene worse, the problem is structural. Change the curriculum, not the clock.
The takeaway
Curriculum learning works because skills compound: master the empty lot before the highway. For an indoor drone, corridors and open halls are the empty lot, and the multi-room apartment is rush hour. The doorway is the single hardest thing it has to do, and getting there requires the right order of lessons, not just more of them.
For the full field report behind this — including the per-scene results that drew the difficulty landscape — see dev-log 49.