What this is: A field report from running the A.1 training curriculum across seven different indoor scenes (S1–S7) and watching which ones our drone learned to fly cleanly — and which one fought back.
Why it’s here: Because the result surprised me. Most scenes were solved quickly, but one scene — a multi-room apartment — got worse the longer I trained on it. That’s a counter-intuitive lesson worth writing down, and it points at exactly where indoor autonomy gets hard: doorways.
Date: 2026-06-15 Ticket: A.1 (FlightRL-v3)
Glossary
A few plain-English definitions before we dig in. If you already live in this world, skip ahead.
-
Curriculum learning — Teaching the agent the way you’d teach a person: easy lessons first, hard ones later. You don’t hand a beginner a maze; you start with a straight hallway, then add a turn, then add rooms. The hope is that skills learned on simple scenes transfer and make the hard scenes learnable. Same idea as a school curriculum — hence the name.
-
Scene — One specific indoor environment the drone flies in. Think of each scene as a different “level” in a video game: a straight corridor, an L-shaped corridor, a two-room flat, a big open hall with columns, a full apartment, and so on. Each scene has its own floor plan and its own coverage challenge.
-
Difficulty landscape — The mental picture of how hard each scene is relative to the others. Imagine a map where flat ground means “easy to learn” and steep cliffs mean “the agent keeps failing here.” Across the seven scenes, the landscape was mostly gentle plains with one stubborn mountain (the apartment). Mapping that landscape tells me where to spend effort.
-
PPO (Proximal Policy Optimization) — The reinforcement-learning algorithm doing the training. In everyday terms: the drone tries things, gets rewarded when it does well, and PPO nudges its behaviour toward what worked — but only nudges a little at a time, never lurching, so learning stays stable. It’s the “learn by trial and error, but cautiously” method.
-
Success / collision rate — How we score a run. Success means the drone completed the coverage mission; collision means it hit a wall or doorway frame. Both are reported as percentages over a fixed set of evaluation episodes (here, 20 deterministic episodes per checkpoint).
1. What I wanted
The goal for this sprint was simple to state: take the minimal coverage reward I’d settled on in the previous dev-log and run a curriculum across the seven A.1 scenes (worlds_a1). I wanted to know, scene by scene, whether the environment, the reward, and the training pipeline actually hold up — and to draw the difficulty landscape: which scenes are trivial, which are genuinely hard.
Mental image: I’m a flight instructor with seven training rooms. I want to know which rooms my student can ace in an afternoon and which ones need a real lesson plan.
2. What I tried
I trained on each scene and evaluated at fixed checkpoints — 100k steps of experience first, and where a scene needed more, 500k. Evaluation was always deterministic over 20 episodes, so the numbers are apples-to-apples across scenes. No reward tweaking between scenes — I deliberately kept the reward identical so that any difference in outcome was about the scene, not about me quietly helping the agent.
The seven scenes ranged from a single straight corridor up to a full multi-room apartment with internal doorways.
3. What happened
Five scenes were solved cleanly. One needed more steps but got there. And the apartment — the one with multiple rooms and real doorways — became the hard case.
Here are the results (minimal reward, deterministic eval over 20 episodes; the “4-4” style notes are coverage waypoints reached):
| Scene | 100k | 500k | Verdict |
|---|---|---|---|
| S1 corridor_straight | ✅ 100%/0% 4-4 | — | solved |
| S2 corridor_L | ✅ 100%/0% 4-4 | — | solved |
| S3 two_rooms | ✅ 100%/0% 4-4 | — | solved (1 doorway) |
| S6 crashtest_zigzag | ✅ 100%/0% 6-6 | — | solved (the crash-test scene!) |
| S5 open_hall_columns | 0%/0% 2-4 (timeout) | ✅ 100%/0% 4-4 | solved at 500k (big hall) |
| S4 apartment | 0%/0% 1-4 (timeout, 0 crashes) | ⚠ 20%/80% crash 1.9-4 | HARD — doorways |
The pattern reads cleanly: single-path scenes (the corridors, the zigzag) are easy. The big open hall just needed more flying time to learn to cover all the area. The apartment is a different animal.
The apartment regression (the interesting part)
This is the bit I want to remember. At 100k steps the apartment policy scored 0% success — but zero crashes. It was timing out: cautious, never reaching the far waypoints, but never hitting anything either. Safe and useless.
At 500k steps it scored 20% success — but 80% collisions. Pure scaling of training steps made it worse in the way that matters. With more experience the policy learned to chase the distant waypoints aggressively, and the place it crashed was the doorways (openings around 1.4m wide, across three to four rooms).
Picture it like a delivery driver who, given more practice, decides to take corners faster to save time — and starts clipping the doorframes. More steps traded safety for aggression. Doorways are the choke point where a bare “get-closer-to-the-waypoint” signal pulls the drone straight into the wall instead of threading the gap.
So the difficulty landscape has a clear shape: single-path scenes are gentle plains; multi-room scenes with doorways are the mountain. The mechanism behind that — that a single corridor is far easier than a multi-room layout with doorways — was confirmed by this run, not just assumed.
Diagnosis / where the apartment goes next
The honest read: this is not a “just train longer” problem. More steps already made it worse. The apartment needs targeted help to get through doorways. Conceptually, the candidates I’m weighing (all cheap, all within our existing rules):
- Curriculum warm-start — start the apartment policy from the already-solved two_rooms scene (which has one doorway it handles fine) and let it build up to many doorways, rather than learning the apartment cold.
- A gentle smoothness/effort signal — on the stable scenes I can afford to encourage smoother motion, which damps the jerky lunges into doorframes. Conceptually: ask the drone to fly cleanly through the gap, not just toward the goal.
- Speed restraint near walls, with a doorway exception — a fallback if the first two aren’t enough: hold the drone back near walls, but carve out an exception so it can still commit to passing through an opening.
The framing that keeps proving true: “fly cleanly through the passage.” Doorways are where that matters most.
4. Sources
- The A.1 scene set (worlds_a1) and the minimal coverage reward established in the previous dev-log.
- Saved tracks viewable in the Player:
a1_openhall_trained— 569 steps, full perimeter sweep, 0.72 mapped. A clean large-area reconnaissance success.a1_apartment_trained— 88 steps, 0.22 mapped. The doorway crash, shown vividly — a short flight that ends at the frame.a1_s1_trained— the straight corridor, mission complete.
- Simulation stack: ROS2 Jazzy + Gazebo + SITL.
5. What’s next
The sprint verdict for the environment and baseline: six of the seven scenes are taken (S1–S3, S5, S6, plus S7 handled separately), with the apartment left as the open hard case — and it’s hard for a specific, understandable reason: doorways. The environment, the reward, and the pipeline are all valid; nothing about the framework is broken.
From here: fix the apartment (warm-start or polish along the lines above), then work toward a single generalist policy trained across rotating scenes, bring S7 (base_stand) into the fold, and check stability across multiple random seeds. The apartment is the lesson; the generalist is the destination.