What this is: a sanity check on last log’s perfect run. We took the exact training recipe that produced one flawless model and re-ran it from scratch with three more random seeds, then re-examined all four models across all 7 training scenes. Why it’s here: in machine learning a single perfect run can be a fluke — the random dice at the start of training just landed well. Before promoting scan-v1 to be the production model of sprint A.1, Aleks asked the obvious question: does it repeat? This log is the answer.
Date: 2026-06-16 Ticket: scan-v1 multi-seed stability (S1/S7), machine D2 (RTX 5070)
Glossary
- PPO — Proximal Policy Optimization, the reinforcement-learning algorithm we train the drone’s “brain” with. It nudges the policy toward better actions in small, safe steps rather than big risky jumps — like adjusting a recipe a pinch at a time instead of dumping in a whole new spice.
- Coverage — how much of the indoor space the drone actually scans during a mission. The core score: did it see what it was sent to see?
- Scene — one test map / room layout the drone has to operate in. We have 7 training scenes.
- scan-v1 — the candidate model line being tested here, named for its scanning behavior.
- A.1 — the current sprint. “Production model of A.1” means the model we actually ship, not a draft.
- Seed — a single number that fixes all the “randomness” in a training run (where the policy starts, the order of samples, etc.). Change the seed and you get a different training trajectory. Same recipe, different dice.
- Multi-seed test — running the same recipe under several seeds. If the result holds across all of them, it wasn’t luck. It’s like checking the same recipe works on three different stoves, not just your own.
- md5 — a short “fingerprint” of a file. Two files with the same fingerprint are byte-for-byte identical; different fingerprints mean genuinely different files.
- Generalist base — a pre-trained starting model shared by all the runs. Each seed fine-tunes from this same foundation.
- Production model — the one that goes into real use, as opposed to an experiment or a throwaway draft.
1. What I wanted
In the previous log (dev-log 52) a single trained model flew through all 7 training scenes perfectly: 100% of tasks completed, zero collisions. Beautiful — but it was one training run.
One great run proves the recipe can produce a great model. It does not prove the recipe reliably produces one. The starting state of training is random, and sometimes the dice just fall your way. So the question I wanted to settle was simple:
If I rerun the exact same recipe with different random seeds, does the perfect result come back — or was scan-v1 a one-off?
If it reproduces, scan-v1 earns the title of production model for sprint A.1. If it doesn’t, we caught a fluke before shipping it, which is exactly the kind of thing you want to catch early.
2. What I tried
I took the identical training recipe — no tuning, no changes — and re-ran it three more times from scratch with different random seeds (seed 2, 3, 4). Each run took about 8 minutes on D2.
That gave me four models total: the original from dev-log 52, plus the three new seeds. Then I put every model through every scene — 4 models × 7 scenes = 28 separate exams.
The deliberate part here is that the only thing changing between runs is the seed. Same data, same hyperparameters, same generalist base to fine-tune from. If something differs in the outcome, it can only come from the random initialization — which is precisely what a multi-seed test is meant to isolate.
3. What happened
All 4 models, on all 7 scenes: 100% of tasks completed, 0 collisions. Every single one of the 28 exams was a clean pass.
| Model (seed) | Scenes passed | Tasks completed | Collisions |
|---|---|---|---|
| original (dev-log 52) | 7 / 7 | 100% | 0 |
| seed 2 | 7 / 7 | 100% | 0 |
| seed 3 | 7 / 7 | 100% | 0 |
| seed 4 | 7 / 7 | 100% | 0 |
So the success was not luck. The recipe reproduces reliably.
A curious detail (and why it isn’t a bug)
The four models flew the routes identically — right down to the same number of steps on each scene. At first glance that’s alarming: maybe the seeds never took effect and I just trained the same model four times, which would make this whole test empty.
So I checked the file fingerprints (md5) of the model weights. They are different. The models are genuinely distinct — the seeds did their job.
The explanation is that all four fine-tune from the same shared generalist base, and for these scenes the optimal flight path to the target points is essentially unique — so everyone converges to the same route. Where the models actually differ is in their scanner behavior: how they sweep the sensor beam around as they fly, not where they fly. Same destination, different way of looking around on the way there.
To be safe I also confirmed the new model’s scanning skill is intact — it actively probes the space ahead of it rather than flying blind. The capability is real, not an artifact of a frozen route.
4. Sources
- dev-log 52 — the original single-seed run that produced the perfect scan-v1 model.
- This log’s 28-exam matrix: 4 seeds × 7 scenes, plus the md5 fingerprint check and the scanner-behavior sanity check.
- The shared generalist base model all four runs fine-tuned from.
5. What’s next
With reproducibility settled, scan-v1 is declared the production model of sprint A.1. From here the work is finishing touches rather than new capability:
- A small polish pass on flight smoothness — the route is right, now make the motion cleaner.
- A check on the “live” Gazebo simulator, a step closer to real hardware than the fast training environment.
The headline takeaway: this was a notable step from “one promising result” to “a recipe we trust.” That’s the difference between a lab curiosity and something you can build the next sprint on.