What this is: A change to how the drone records its understanding of a space. Instead of saving one flat snapshot of “semantic” map data, the flight-log now carries a living, timestamped stream of how the drone colors zones by meaning as it flies.
Why it’s here: Aleks asked for something specific — the drone paints the map by meaning during flight, and that coloring should be saved into the flight-log with timestamps, so the Player can show it appearing and growing over time. This entry documents how I rebuilt the semantic record to make that possible.
Date: 2026-06-16 Ticket: rl-lab dev-log 38
Glossary
- Dynamic semantic stream — a running feed of “what the drone now understands about each part of the map,” colored by meaning, saved continuously instead of as one fixed picture. Think of it like watching someone slowly color in a coloring book over a video, rather than seeing only the finished page.
- Flight-log — the recording of a flight. It already stores things like where the drone went and what its sensors saw; now it also stores the semantic coloring as it changes. Imagine a flight recorder (“black box”) that also remembers what the pilot figured out at each second.
- Delta — we only write what changed since last time, not the whole map again. Like saving only the edits to a document instead of a fresh copy every keystroke.
- Keyframe — a full snapshot taken every few seconds, so a viewer can jump to any moment without replaying everything from the start. Like chapter markers on a video.
- Sink — the destination where records flow to (here, the flight-log file). Think of it as the drain that all the writes pour into.
- Info-gain — “how much new the drone learned.” Reward is tied to learning new ground, not to how many kilometers were flown.
- PPO — the reinforcement learning algorithm that trains the flight policy. It is the brain being trained; this change feeds it nothing new — the semantic stream is a separate recording channel.
1. What I wanted
The goal Aleks set is easy to picture: the drone should color the map by meaning while it flies — “this region is open and safe to go,” “this is mapped,” “this is a task area” — and that coloring should land in the flight-log tagged with the moment it happened.
The payoff is in the Player. When you scrub the timeline, you should literally watch the colored zones bloom and spread. This matters most for simulation flights where Aleks flies manually: a human pilot already knows what to color and why, so the simulation has to do that coloring on the drone’s behalf, automatically, the same way the real policy would.
2. What I tried
The flight-log format already had a “semantic” line, but it was nearly empty — a single layer, usually just one frame written at the very beginning. A dead photograph instead of a movie. I rewrote it into a live coloring stream:
- Three layers per cell — navigator, cartographer, and tasks. Each map cell carries all three meanings independently, so one cell can be “safe to fly,” “mapped,” and “task-relevant” at once.
- Delta writes — we only record what changed since the previous tick. Without this, the file would balloon: re-saving the full map every frame would bloat the log enormously.
- Keyframe every ~5 seconds — a full snapshot at a fixed interval so the Player can jump anywhere on the timeline without replaying from frame zero. The deltas fill in the moments between keyframes.
- Colors from the shared zone config — the same configuration I set up the day before, so the Player and the recorder agree on what each color means. No duplicated, drifting color tables.
3. The main rule (anti-vacuum-cleaner)
This is the rule that shapes everything: a cell gets colored when the drone has SEEN it with a rangefinder beam — not when it has flown into it. It is enough to understand “you can go there, it’s safe.” The drone does not need to actually visit the cell.
Aleks has said this many times: the drone is not a robot vacuum cleaner. So the reward for “uncovered a new part of the map” is also given for what the drone saw, not for what it physically circled. If it loops around already-seen ground, there is no reward for that — in fact, there’s a light penalty for treading the same floor.
The mental image: the semantic stream in the flight-log is a map of what the drone knew at each moment. That is exactly the thing you see in the Player, growing along the timeline. Info-gain — how much new it learned — is the currency, not distance traveled.
4. What happened
The semantic record went from a single near-empty frame to a continuous, scrubbable stream. The three-layer, delta-plus-keyframe structure keeps the file size reasonable while giving the Player everything it needs to render the coloring at any chosen moment: accumulate the deltas up to that instant, then paint by the shared config.
Importantly, this does not touch the model’s inputs or outputs. The PPO policy sees nothing new — the coloring rides as a separate stream into the flight-log. It is a recording feature, not a training change.
5. Sources
- The flight-log format and its existing “semantic” line (the single-frame version I replaced).
- The shared zone color config built the previous day, reused here so the recorder and Player stay in sync.
- The existing scan-point recording path in simulation, which the new semantic stream mirrors.
6. What’s next / who I handed off to
- To the interface: add Player support for showing these changing zones over time — accumulate changes up to the selected moment, color by the config, and offer two display modes.
- To simulation: their module that decides zones should send its coloring (driven by beam visibility) into the same flight-log, exactly as already done for scan points.
The model’s input and output are unchanged by any of this — the coloring travels as a separate stream into the recording.