What this is: A tooling check, not a training run. I took a real drone flight-log and pushed it through the entire pipeline — record it, replay it, hand it to the mapper — to make sure all three pieces actually fit together.
Why it’s here: Three different teams each built one link of this chain and tested their link against made-up sample data. Nobody had ever run real data of our drone through the whole thing at once. I did, and I closed two open format questions while I was at it.
Date: 2026-06-14 Ticket: 36 — track pipeline consolidation test (no training, tooling prep)
Glossary
Plain-English definitions for the moving parts in this entry.
- Track pipeline. Think of it as a relay race with three runners. The drone flies and writes down everything it sees and does (a flight-log). That log can later be replayed like a video, and it can be handed to a mapper that turns scanned points into a map. The “pipeline” is the whole baton-passing chain: fly → record → replay → map.
- Consolidation. Bringing the separate links together and checking that the baton actually passes cleanly from one runner to the next. Each team had only tested its own leg in isolation; consolidation is the first time the full relay runs end to end on real data.
- Regression test. A safety net. The idea is simple: take something that already worked before (an old, known-good flight-log), run it through the new tools, and confirm it still works. It guards against the classic trap where fixing or changing one thing quietly breaks another.
- PPO. The reinforcement-learning algorithm we use to train the drone’s flying policy. It learns by trying actions, seeing how well they paid off, and nudging its behavior toward what worked — but cautiously, never lurching too far from what it already knew. No training happened in this entry; PPO is just the consumer this pipeline ultimately feeds.
A few more terms specific to our stack:
- flight-log. A line-by-line record: where the drone was, what it commanded, what it scanned. Replayable like a video.
- gen_0 / “raw”. Sensor points exactly as they arrived, never edited afterward.
- mapper. The program that builds a map from scanned points and learns to correct drift — the gap between where the drone thought it arrived and where it really was.
- bearing. The angle of the scanner’s beam — which direction the sensor was looking.
- Map Studio. Our web-based flight replayer with a time-scrubber, so you can rewind and step through a flight frame by frame.
1. What I wanted
We’re building a pipeline around drone flights, and the shape of it is this:
- The drone flies a mission.
- A flight-log gets written, line by line.
- That log can later be replayed like a video (scrub back and forth in time).
- The same log can be handed to the mapper, which turns scanned points into a map.
Three teams own three links: interface owns the log format and the replayer, simulation owns the scanned points coming off the sensor, and I own the format rules plus the drone itself. Up to now, each of us had checked our own link against invented test data. That’s fine for catching obvious mistakes, but it tells you nothing about whether the batons actually pass between runners.
My goal for this entry was the boring-but-important one: run real flight data from our own drone through the entire chain, in one go, and see whether it survives intact. This is a regression test in spirit — prove the new tooling doesn’t quietly mangle data that used to be fine.
While I was in there, I also wanted to settle two design questions that had been hanging over the format.
2. What I tried
I attacked it from two angles — one looking backward, one looking forward.
The two open questions first. These were holding up the format spec, so I made the calls:
- Should the mapper store “range” (how many meters to a point) as its own dedicated column? — Yes. Range is the sensor’s honest, original reading; it doesn’t change just because a human later drags a point around on the map by hand. If you instead recompute range backward from a point that’s been nudged, you’ve corrupted the original measurement. So we keep range as a separate column, computed from the raw (gen_0) reading. Picture it like a receipt: even if you later reorganize your budget spreadsheet, you don’t go back and rewrite what the store actually charged you.
- Does the drone need to scan while moving (in flight) right now? — No. Today the drone scans while stationary: it hovers, looks around, then flies on. Scanning on the move is a harder skill, and we deliberately deferred it. Keeping the drone still during a scan keeps the model simpler, which matters for the PPO policy that eventually has to learn this behavior.
Then the actual end-to-end runs. Two of them:
- The backward-looking run: I took an old, real flight-log — 884 rows from a previous SITL run — and opened it in the new replayer. The point was to make sure last month’s known-good data still flows through this month’s tools.
- The forward-looking run: I recorded a fresh log in the new format, read it straight back, and handed it to the mapper, to confirm a clean round-trip on data born in the new world.
3. What happened
Both runs passed, and the details lined up cleanly.
| Run | Source | Rows | Result |
|---|---|---|---|
| Backward (regression) | Old real flight-log, prior SITL run | 884 | Opens in the new replayer; angles convert correctly (degrees → radians); every step carries both “where the drone thought it was” and “where it actually was” |
| Forward (round-trip) | Fresh log, new format | — | Recorded, read back, handed to mapper; everything present; range restored exactly: [1.5, 2.0, 1.7] m |
A couple of things I want to call out, because they’re the parts that could quietly have gone wrong:
- Angle units. The old log stored angles in degrees; the new replayer expects radians. The conversion came out right on every step, which means we won’t get a flight that visually “looks fine” but is silently rotated.
- The two positions stay paired. Every step keeps both the drone’s believed position and its true position. That pairing is exactly what the mapper needs later to learn drift correction — if it had dropped one of the two, the whole point of the log would evaporate.
- Range survived the round-trip exactly. The mapper got back [1.5, 2.0, 1.7] m, matching the raw readings to the digit. That’s the strongest evidence the “store range as its own column” decision works as intended.
I also produced a ready-made sample file so Aleks can open it in Map Studio and eyeball the replay himself. I deliberately did not launch the heavy browser session on my side — I’m keeping memory free, and a human’s eyes on the replay are a better check anyway than me describing pixels.
4. Sources
- Old real flight-log (884 rows) from a prior SITL run — the regression input.
- Fresh new-format log recorded during this entry — the round-trip input.
- The new replayer and the mapper consumer, owned across the interface and simulation teams.
- ROS2 (Jazzy) and Gazebo as the surrounding simulation stack; SITL as the run mode that produced the original log.
5. What’s next
The pipeline is proven on real data, so the immediate follow-ups are about turning that into routine:
- Hand the sample file to Aleks for a visual pass in Map Studio.
- Treat this end-to-end run as a repeatable regression check, so future format changes have to clear the same “old log still flows, new log round-trips” bar before they land.
- With scanning-in-flight explicitly deferred, the stationary-scan assumption stays baked into the format — which keeps the PPO training setup simpler for now.