claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

45 — Loading the A.1 Worlds Into the RL Environment

The RL drone moves out of an empty test box and into the seven real A.1 sprint scenes — corridors, an apartment, a pillar hall — with maps loaded and verified.

stablerl-labupdated 2026-06-16T00:00:00.000ZClaudeDroneRLDevLog

What this is: the log of the day the reinforcement-learning drone finally stopped flying inside a featureless square test room and started flying inside the actual A.1 sprint worlds — real corridors, a multi-room apartment, a pillared hall, a zigzag stress course.

Why it’s here: training a drone in a toy box teaches it almost nothing about the spaces it has to cover for real. This entry marks the moment the RL environment can ingest any of the real scenes the simulation team built, read them correctly, and play back flights in the right coordinate frame. It’s the plumbing step that everything downstream depends on.

Date: 2026-06-15 Ticket: rl-lab dev-log 45 — load A.1 worlds (S1–S7) into the RL environment


Glossary

  • Occupancy map — a grid that records, cell by cell, whether a spot is wall, free space, or simply unknown. Picture graph paper laid over a floor plan, where each square is colored in once you know what’s there. The drone navigates against this grid the way you’d read a maze on paper before walking it.
  • A.1 worlds / scenes (S1–S7) — the seven real test environments for this sprint, authored by the simulation team. They run the gamut from a tight corridor to an L-shaped room, a full apartment with separate rooms, a hall studded with pillars, and a long zigzag crash-test course. Think of them as a set of increasingly nasty obstacle courses rather than one empty gym.
  • Spawn point — the exact spot where the drone starts each run. Like the “Start” tile on a board game; if you put the piece on the wrong square, everything after it is off.
  • Track (flight log) — a recording of a single flight, saved so our Player can replay it later. It’s the flight’s black box: where the drone went, what its sensors saw, frame by frame.
  • .npz map file — the packaged map format the simulator exports. One file holds the occupancy grid plus metadata (size, scale, spawn). Treat it as a sealed box that the environment has to unwrap and trust.
  • Player / Map Studio — the visualization tool where you scrub through a recorded flight and watch the drone move over the map. Like a video player, but for flights.
  • Policy (“the brain”) — the trained decision-maker that, eventually, tells the drone where to go. Until it’s trained, the drone is effectively a learner driver with no instructor: it just guesses.

1. What I wanted

Up to now the RL drone lived in a single small empty square room, 6.4 × 6.4 m, with nothing in it. That’s a toy. A policy trained there learns to fly around an empty box, which is roughly as useful as practicing parallel parking in an empty parking lot the size of a tennis court — technically driving, but not the skill you actually need.

The goal for this entry was simple to state and fiddly to deliver: make the RL environment able to load the real A.1 sprint worlds. Not a re-creation of them, not an approximation — the very same maps the simulation team produces, read straight from their .npz exports, whatever shape and size they happen to be.

2. What I tried

Three pieces of plumbing, each one a place where “close enough” would have quietly poisoned every future training run:

  1. Teach the environment to read simulator maps of any size. The old code silently assumed a 64 × 64 square grid. The real worlds are nothing like square — a corridor is long and thin (10 × 2 m), the apartment is a big 12 × 10 m floor plan carved into rooms. So I made the loader accept the occupancy grid at whatever dimensions it arrives in. While I was there I fixed the semantics: a cell that’s wall and a cell that’s “unknown” both count as something-the-drone-must-not-fly-through. Unknown isn’t free; treating it as free is how you get a drone that confidently flies into a fog of walls.

  2. Take the spawn point from the map itself. Instead of dropping the drone at some hardcoded corner, the environment now reads the proper start position out of the map, exactly the way the full simulator does. This matters more than it sounds: if the spawn drifts even a little, the drone can start half-inside a wall, and then nothing downstream makes sense.

  3. Record real flights and wire them to the Player. I flew runs in two of the new worlds — the corridor and the apartment — and saved the tracks so they open in Map Studio. Then I checked that the playback actually lines up: the drone’s coordinates match world coordinates, and the individual scan points land cleanly on the walls instead of floating beside them.

3. What happened

It loaded, and — the part that actually let me exhale — it loaded correctly. I cross-checked every one of the seven worlds against the reference numbers the simulation team sent over: the physical dimensions and the count of free (navigable) cells. They matched down to the cell. Not “roughly the same area,” not “within a few percent” — exact agreement on the free-space count, which is the strict test, because a single off-by-one in how you read a grid usually shows up as dozens of misplaced cells.

Check Result
Worlds loaded (S1–S7) all 7 read successfully
Non-square maps (corridor 10×2, apartment 12×10, etc.) handled at native size
Dimensions vs. simulation reference match
Free-cell count vs. reference match, cell-for-cell
Spawn point read from map, matches simulator
Recorded flights corridor + apartment, replay in Player
Scan points vs. walls in playback aligned, sit on walls

One thing worth saying plainly so nobody misreads the recordings: in these tracks the drone flies more or less at random and crashes into things. That is expected and is not a sign that the physics or the maps are broken. The drone simply has no trained brain yet — no policy. It’s a student driver with no instructor and no map memorized, so it lurches around and hits walls. The point of this milestone was the maps and the coordinate plumbing, both of which are now solid. The flying-well part comes after training.

4. Where to look

The tracks live under the interface save area for the RL flights:

  • interface/save/rl/a1_corridor/
  • interface/save/rl/a1_apartment/

Open the .flight.jsonl file in Map Studio (the Player) to scrub through either run and watch the drone move over the real map, with scan points landing on the walls.

5. Sources

  • The seven A.1 sprint worlds (S1–S7) and their reference dimensions and free-cell counts, provided by the simulation team.
  • Simulator .npz map exports, read directly by the RL environment.
  • Recorded flight tracks in the corridor and apartment worlds, replayed in Map Studio.

6. What’s next

The environment can now place the drone in real spaces, but it’s still flying blind and untrained. Two things follow naturally from here: adding a richer, color-coded semantic layer on top of the bare occupancy grid so the drone perceives more than just wall-versus-free, and then actually training a policy so the flying stops being random and starts being coverage. The maps are no longer the unknown variable — that’s the quiet win of this entry.

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR