What this is: We packaged the trained SWEEP-02 policy into a self-contained folder,
export/sweep02/, so the simulation agent can plug it into Gazebo. The pack holds seven files: the model itself, a written description of its inputs and outputs, a lightweight Python stand-in for the simulator, and two automated tests. Why it’s here: This is the very first step of the model-to-sim handoff. Without this pack, the simulation team has nothing concrete to build a ROS2 bridge against. This entry shows what a clean model handoff looks like in practice.
Date: 2026-05-19 Ticket: TASK-058 (model-to-simulation sprint, Phase 1)
Glossary
A quick cheat sheet so the rest of the entry reads smoothly.
- Export pack — a single folder that travels from the training team to another team. Think of it like a shipping crate: it contains the finished product (the trained model) plus the manual (input/output specs) and a couple of inspection tools so the receiver can confirm nothing broke in transit.
- Model handoff — the moment one team gives a trained model to another team to use. Like handing over a car: the new driver needs the keys (the model file) and the owner’s manual (how to talk to it), or they’ll get nowhere.
- SWEEP-02 model — our best-trained network for the 2D scenario. It decides where the drone should move and turn next.
- Bridge — the connector between Gazebo and our model. It takes sensor readings, asks the model what to do, and sends the resulting movement command back to the drone.
- Observation — the data fed into the model: distances to walls from the drone’s sensors, the current servo angle, and a small grid recording where the drone has already been.
- Action — the model’s output: a single decision such as “go forward,” “turn left,” or “fly toward the wall.”
- Gazebo — a 3D physics simulator where the drone behaves almost like a real one.
- Mock environment — a “fake Gazebo” written in plain Python. It lets us test the bridge without launching the heavy 3D simulator every time.
- Coverage — the share of the room’s free space the drone actually visited. This is our headline metric; a coverage of 0.62 means the drone reached 62% of the open floor.
- Inference — the instant the model “thinks”: it receives one observation and returns one action. This needs to be fast.
1. The goal
The sprint is called model-to-sim-bridge: take the model we trained in a 2D world and connect it to the Gazebo 3D simulator. For that to work, the simulation agent needs to know exactly what we send the model and exactly what it sends back. Even a small mismatch in the data format and the model will behave strangely or simply refuse to act.
My slice of the work (TASK-058) was to assemble a “model passport” plus tests in the export/sweep02/ folder. Nothing more, nothing less — I do not write the bridge itself, and I do not launch Gazebo. That belongs to the simulation team.
2. What went into the pack
Seven files, each with a clear job:
model.zip(4.2 MB) — the trained model itself. This is not source code; it is the network’s learned weights, packaged in the standard format produced by our PPO training stack.obs_spec.md— a written description of the model’s input. It explains which sensors are used, how each reading is normalised (for example, a wall distance in metres is rescaled into a tidy number between 0 and 1), and how the “where I’ve already been” grid is laid out.action_spec.md— a written description of the output. There is one action per decision, and each carries its own meaning and duration. One detail matters a lot: the “fly forward to the wall” action must be understood by the bridge as a continuous motion that runs until the drone reaches the wall, not a single one-off nudge. Getting this wrong is the most likely way to confuse the model.model_card.md— the model passport. It records how the model was trained, the coverage it reached at different training lengths on the 2D scenario, the targets we expect to hit in Gazebo, and a clear retrain trigger: if Gazebo coverage falls below a stated floor, the model should be retrained in an environment tied to Gazebo physics.mock_env.py— the fake Gazebo in Python. The simulation agent can exercise the bridge against this stand-in instead of spinning up the full simulator for every test.validate_export.py— an automated correctness test. It runs the model through the mock environment many times and checks that coverage still lands near the figure we measured before packaging. If it drifts, something broke during export.inference_benchmark.py— a speed test. It measures how long the model takes to “think” about a single observation.
Here is the shape of the pack as it sits on disk:
export/sweep02/
├── model.zip # the trained policy (weights)
├── obs_spec.md # what goes in
├── action_spec.md # what comes out
├── model_card.md # the model passport
├── mock_env.py # fake Gazebo for testing
├── validate_export.py # correctness test
└── inference_benchmark.py # speed test
3. What came out
Correctness test (validate_export.py)
I ran 100 episodes across five evaluation maps:
| Measure | Result |
|---|---|
| Mean coverage | 0.6174 |
| Calibration reference (before packaging) | 0.6186 |
| Difference | 0.001 — well inside random noise |
| Minimum episode | 0.532 |
| Maximum episode | 0.698 |
| Easy map (map_01) | 0.65 |
| Hard map (map_02) | 0.58 |
The numbers match the reference almost exactly, and the easy/hard split goes the way it should — easier rooms get cleared more thoroughly. Verdict: green. The model was packaged correctly.
Speed test (inference_benchmark.py)
I called the model 1000 times in a row on a plain CPU, with no graphics card:
| Measure | Result |
|---|---|
| Mean think time | 0.34–0.70 ms |
| 99% of calls under | 0.75 ms |
| Worst call observed | 0.78 ms |
| Target | under 10 ms |
We came in roughly 20× faster than the target. In practice this means the bridge can comfortably run at 10 Hz — one decision every 100 ms — with plenty of headroom left over for reading sensors, sending movement commands, and logging.
The mock environment behaves
Running the policy through mock_env for 100 episodes produced essentially the same coverage as the real training environment (0.617 versus 0.619). The mock simulates the physics faithfully enough that the simulation agent can trust it for early bridge development — no need to wake up Gazebo just to check that the wiring is correct.
4. Useful links
- Sprint plan v2 (my notes on the input/output specs are folded in):
_workspace/agents/orch/sprint_model_to_sim_v2_final.md - Ticket TASK-058:
_workspace/tickets/TASK-058.md - The canonical experiment that
model.zipcame from:experiments/TASK-RL-SWEEP-02-seed22/ - The earlier write-up describing the model itself:
docs/dev-log/14-sweep-and-potential-null.rl-lab.md
5. What happens next
- Right now — send the handoff note to the simulation team with links to all seven artifacts, and log it in the shared sprint journal.
- This session — the simulation agent starts writing the ROS2 bridge node (ROS2 Jazzy), reading my specs as they go. Technical questions can come straight to me while the sprint is live.
- A few days out — a joint checkpoint: once the simulation team’s mock tests pass, we run the bridge against my
mock_envtogether and confirm the coverage still lands at 0.6174. - In parallel — a room layout spec (TASK-060): I describe three or more calibration rooms (empty, one column, corridor) and the simulation team builds them as SDF worlds for Gazebo.
- Then — Aleks, the simulation team, and I calibrate the physical movement constants in Gazebo, run a full episode, and measure the gap between simulation and reality.
What this task is not
- I do not write the bridge node — that is the simulation agent’s job.
- I do not launch Gazebo — at this stage the training side only trains models; Gazebo testing happens through the simulation team.
- I do not retrain the model — SWEEP-02 is already our production policy. Retraining only happens if Gazebo coverage drops below the agreed floor, and even then it would be done in an environment tied to Gazebo’s physics.