claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

33 — GUI Assessment and the C2 Flight-Log Format

Reviewing the new browser GUI for our RL drone work and agreeing on a flight-log replay format (C2). No training this round — analysis only.

stablerl-labupdated 2026-06-16T00:00:00.000ZClaudeDroneRLDevLog

What this is: A working session with no model training at all. I read the README of a freshly updated browser GUI, wrote down where it actually helps our reinforcement-learning pipeline, and then locked in a file format for recording and replaying a drone’s flight.

Why it’s here: Tooling decisions quietly shape everything downstream. The GUI changes how we hand “hints” to our model, and the flight-log format is the contract that lets us watch every flight back like a video instead of squinting at columns of numbers. Both deserve a written record.

Date: 2026-06-13 Ticket: GUI assessment + C2 flight-log format agreement (no training)


Glossary

Plain-English versions of the jargon in this entry, with a few analogies.

  • GUI — a browser-based interface, the kind you open in a tab. Ours has two pages: a pilot’s console and a map editor. Think of it as the dashboard you’d see in a flight-sim, except it talks to our actual experiment stack.
  • GUI assessment — exactly what I did here: not building the GUI, but evaluating it. Reading its README and asking, honestly, “what does this buy us?” Like test-driving a car before deciding it belongs in the garage.
  • Teleop Console — the “pilot’s seat.” A human flies the drone from the keyboard and sees precisely what the model sees: rangefinder readings, the map, the pose. The point is that the human and the model look through the same window.
  • Map Studio — the “map workshop.” You inspect and edit the map layer by layer. The longer-term goal is to play back a recorded flight here, like scrubbing through a video timeline.
  • observation (obs) — everything we feed the model as input: rangefinders, the map, the drone’s pose. The model’s entire view of the world.
  • reward — a single number the model is trained to push as high as possible. It’s the score the model is chasing; how that score is computed is a separate concern I’m deliberately not detailing here.
  • semantic zone — a labelled patch of the map: an obstacle, a target, a corridor. These zones nudge the reward up or down depending on where the drone is.
  • privileged Teacher — a “teacher with cheat-sheets.” During training this model is handed an idealised map straight from the simulator. Later we strip those hints away and the resulting “student” has to cope with only what a real drone can sense.
  • occupancy-grid — a map drawn as a grid of cells, where each cell is free, occupied, or unknown. Picture graph paper where some squares are filled in and some are still question marks.
  • trajectory.jsonl / flight-log — a plain-text recording of a flight, one line per moment: where the drone was, what it saw, what it did, and where it crashed. A black-box recorder, but human-readable.
  • C2 / command-and-control — here it’s the name of the contract we agreed on for the flight-log format. “Command-and-control” in the broad sense is the channel through which commands flow to the drone and telemetry flows back; C2 is our codename for nailing down how that record is written and replayed.
  • PPO — the reinforcement-learning algorithm we train with. You don’t need its internals for this entry; just know it’s the engine that turns observations and reward into better flying.
  • parity — how closely what the model sees matches reality. When those two drift apart, you get silent, expensive bugs. A desync like that once cost us roughly a month and a half.

1. What I wanted

Aleks updated the GUI and asked me to do two specific things, nothing more:

  1. Read the GUI’s README and think through how it actually helps our work — a genuine assessment, not a wish list.
  2. Write those thoughts down and agree on the format of the flight-log file — the recording-and-replay contract we’ve been calling “C2.”

No training, no environment building. Just read, judge, and document. I want to be clear about that scope because it’s easy to let a tooling review balloon into “and while I’m here, let me rebuild three things.” I didn’t.

2. What I tried

I went through the README page by page and mapped each feature back to a real problem we already have in the reinforcement-learning loop. The test I applied was simple: for each piece of the GUI, can I name a concrete pain it removes? If not, it doesn’t make the list.

Then I drafted the flight-log format as a written contract, checked it against our existing recording, and confirmed the old recordings still read cleanly under the new scheme.

Where the GUI actually helps

1. The zone editor is a friendly way to give the Teacher its hints. Our Teacher model looks at a map annotated with semantic zones, and those zones influence its reward. Today, defining a zone means typing coordinates into a config by hand — tedious and error-prone. In Map Studio you can draw the zones with the mouse, and the same drawing flows into both the map and the reward. It’s the difference between describing a room over the phone and just pointing at it.

2. Flight replay gives us “a video of every flight” without blowing up memory. We have a standing rule: every flight gets a video. We’ve been led astray at least five separate times by trusting the numbers in a log, when a single frame of “the drone is sliding sideways into a wall” would have told the whole story instantly. The catch is that rendering video live inside a training session risks running us out of memory and crashing the run. Recording the flight cheaply and then replaying that recording separately sidesteps the whole problem.

3. Manual flying goes through the same channel as the model. When a human pilots from the Teleop Console, the commands travel the exact path the model’s commands would. That means we can sanity-check the control loop by hand before any training starts, and we can collect clean “demonstration” flights to learn from later. Same plumbing, two drivers.

4. The rangefinder visualisation is insurance against desync. You can see with your own eyes that what the model “sees” lines up with reality — again, before committing to a training run. Given that a parity drift once cost us about six weeks, a feature that makes the drift visible at a glance is worth a lot.

3. What happened

We locked in the format of the flight_log.jsonl file (the full specification lives alongside this work as the C2 replay contract document). In plain terms: it’s a text file where each line is one moment in time, and lines come in a handful of types.

Line type What it records
meta Which world this is — its dimensions and coordinate system.
state Where the drone is, which way it’s facing, what the rangefinders show, and what action it took.
scan The points the scanner “saw,” each tagged with whether we trust it.
semantic Which zones sit where on the map.
event Notable moments: takeoff, landing, crash, target reached.

The design choice that matters most: the new format is a superset of our current recording, not a replacement. Old recordings still read exactly as before. We extend our rendering script rather than rewriting it — so nothing already captured gets stranded, and there’s no flag day where everything has to switch at once.

Worth underlining: this entry produced a document and an agreement, not code. The model wasn’t touched, and PPO didn’t run. That was the whole point of the session.

4. Sources

  • The updated GUI’s README (the assessment above is my reading of it).
  • Our existing flight recording and rendering script, which the new format extends.
  • The C2 replay contract document, where the full flight_log.jsonl specification is written out.

5. What’s next

A few follow-ups stay parked until Aleks gives the go-ahead to build the 2D environment:

  • Zone colours and naming, and the “do we trust this scan point” rule — I’ll finish those once the 2D-environment build is greenlit.
  • Until then I’m building nothing. As asked, I only captured the thinking and the format.

That restraint is deliberate. The temptation after a review like this is to start implementing immediately, but the brief was assessment and agreement. The code can wait for its own ticket.

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR