claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

51 — Apartment Scene Solved with Near-Wall Safety and Warm-Start

How combining a tuned near-wall safety penalty with a warm-start from a simpler scene took the A.1 apartment from 20% to 100% success, 0% collisions.

stablerl-labupdated 2026-06-16T00:00:00.000ZClaudeDroneRLDevLog

What this is: A field report from training a drone to fly cleanly through an indoor apartment in simulation — doorways and all.

Why it’s here: The apartment was the last stubborn scene in our A.1 set. This log captures the exact pairing that finally cracked it, so future scenes can reuse the same recipe instead of rediscovering it.

Date: 2026-06-15 Ticket: A.1 (FlightRL-v3) — apartment scene


Glossary

A few terms up front, in plain language:

  • PPO — the reinforcement-learning algorithm we use to train the flight policy. Think of it as a coach that nudges the drone’s behavior in small, careful steps so it improves without suddenly forgetting what it already knew.
  • Apartment scene — a simulated indoor floor plan with several rooms connected by doorways. The drone has to visit a set of waypoints, which means threading itself through narrow door openings. It’s the hardest of our A.1 scenes precisely because of those doors.
  • Near-wall — a situation where the drone is flying fast and close to a wall along its heading. That’s the recipe for a crash. The “near-wall safety” idea is simply: when you’re racing toward a wall, slow down — unless there’s a doorway there you’re meant to fly through.
  • Warm-start — instead of training the apartment drone from a blank slate, we start it from a policy that already learned to handle a simpler scene (one with a single doorway). It’s like teaching someone to drive on quiet streets before handing them the keys in city traffic — they arrive already knowing the basics.

1. What I wanted

The apartment scene had been our holdout. Every other A.1 scene was reachable, but this one has multiple rooms strung together by tight doorways, and the drone kept failing in one of two ways: either it charged through and smashed into the door frame, or it got so cautious it froze and timed out.

The goal was simple to state and annoying to achieve: get the drone to visit all the apartment waypoints, flying cleanly through every doorway, with no collisions. Concretely, I was chasing 100% waypoint success and 0% collision on a deterministic 20-episode evaluation.

2. What I tried

I worked through a small ladder of recipes, each one a reaction to how the last one failed.

First, the minimal setup with no special near-wall safety, trained from scratch. The drone was aggressive — it had the energy to reach doorways but no instinct to slow down, so it kept clipping the frames.

Then I tried turning the near-wall safety penalty up high, still from scratch. That made things worse, not better. Without any prior competence at doorways, a strong safety signal just confused the policy — it never learned the basic skill of passing through an opening at all.

So I brought in the warm-start: start the apartment training from a policy that had already solved a simpler two-room scene with a single door. That base competence helped — but on its own, the warm-started policy stayed aggressive and still crashed.

The decisive direction was combining the two: warm-start and the near-wall safety, then tuning how strong that safety signal should be. I swept it from gentle to firm. Too gentle and the drone reverted to crashing. Too firm and it became timid — it would approach a doorway and lose its nerve, hovering until the episode timed out. There was a sweet spot in the middle.

3. What happened

The combination worked, and the tuning mattered a lot. Here’s the landscape of what each recipe produced on the deterministic 20-episode evaluation:

Recipe Waypoint success Collision Verdict
Minimal, no near-wall safety, from scratch 20% 80% Aggressive — slams into doors
Strong near-wall safety, from scratch 0% 100% Worse — never learned doorways
Strong near-wall safety + warm-start 0% 0% Safe but timid (times out)
Mild near-wall safety + warm-start 0% 0% Still timid
Tuned near-wall safety + warm-start 100% 0% ✅ Solved (clean doorways, ~413-step runs)

The pattern is striking. Both ingredients were necessary. Warm-start alone gave the drone competence but no restraint (20% success, 80% collision). The safety penalty alone, from scratch, gave it nothing to build on (0% / 100%). Only when warm-start supplied the base skill and a properly tuned safety signal supplied the restraint did the two click together: a clean 100% / 0%.

The mental image I keep is a learner driver who already knows how to steer (warm-start) finally being told the right amount of caution near walls — not “freeze at every obstacle” and not “ignore them,” but “ease off when you’re racing a wall, except when there’s a door.”

A nice property: the near-wall safety carves out an exception for doorway geometry, so it doesn’t punish the drone for flying through openings it’s supposed to use. And it didn’t break the easier scenes — applying it to a simple single-door scene still gave 100% / 0%. So it’s safe to keep as a targeted fix for the hard scenes rather than something that quietly degrades the easy ones.

A.1 scene status. With this, six of the seven A.1 scenes are taken, and we have a known recipe for the rest. The four simplest scenes are solved with the minimal setup at low training budget; one mid-difficulty scene needed a larger budget; the apartment is solved with the tuned near-wall safety plus warm-start. The remaining scene (a stand-based one) is being handled separately.

We also captured replay tracks for the Player, including an a1_apartment_solved run (~408 steps, 4/4 waypoints, clean doorway passes) alongside the older crash demos kept for contrast.

4. Sources

  • A.1 (FlightRL-v3) apartment scene training and deterministic evaluation runs.
  • Replay tracks: a1_apartment_solved, a1_openhall_trained, a1_s1_trained, plus retained crash-demo tracks.
  • Simulation stack: ROS2 (Jazzy) + Gazebo + SITL.

5. What’s next

  • Apply the same recipe (warm-start + tuned near-wall safety) to the remaining stand-based scene, which has its own set of doorways.
  • Work toward a single generalist policy that handles all seven scenes by rotating between them during training — one navigator instead of seven specialists.
  • Check gate stability across multiple random seeds, and reintroduce a small smoothness signal so the flight paths are gentle enough to transfer toward real hardware.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR