What this is: A second, independent test of the “bridge” — the piece of code that shuttles a drone’s sensor data into the trained policy and the policy’s actions back out to the simulator — run over full-length episodes instead of short ones. Why it’s here: The simulation team had already tested the bridge, but only with very short runs. Short runs are great at catching obvious crashes and terrible at catching slow, accumulating drift. So we re-checked it the hard way, cell by cell, over thousands of steps. The result is clean: every episode matched perfectly.
Date: 2026-05-19 Ticket: rl-lab bridge cross-side sanity
Glossary
Plain-English definitions first, so the rest of the article reads smoothly.
- Bridge — a small middleman program (the
policy_bridgepackage). Think of it as a translator standing between two people who speak different languages: the simulator (Gazebo) on one side, and our trained policy on the other. It hands each side what it expects and nothing more. - Cross-side sanity check — testing a component by comparing two independent computations of the same thing and confirming they agree. Here, two different programs each build a map of “where the drone has been” and count coverage. If the bridge is correct, both maps must be identical. It’s like two cashiers counting the same till separately and getting the same total — agreement is the proof.
- Bit-exact — not just “close enough,” but identical down to the smallest unit. Every single cell of the grid matches, and the coverage number matches out to five decimal places. There is no rounding, no “within a percent or two.” It’s the difference between two photographs that look alike and two photographs that are the same file.
- Coverage — the fraction of the reachable indoor space the drone has actually flown over. If a room has 100 open floor cells and the drone has visited 95 of them, coverage is 95%.
- Visited grid — a 64×64 checkerboard laid over the room. Each square is either “visited” or “not visited.” As the drone flies, squares flip to “visited.”
- odom — short for odometry. In ROS2 (Jazzy) this is the
/mavros/local_position/odommessage stream — the simulator’s running estimate of where the drone is, broadcast many times per second. - PPO — the family of reinforcement-learning algorithm behind our policy. For this article it’s enough to know it takes in an observation and returns one of a small set of moves.
- SITL — software in the loop: running the flight stack against a simulator instead of real hardware, so we can test behavior safely before flying anything physical.
1. The setup: why test the same thing twice?
When the simulation team finished the bridge, they ran a batch of short test scenarios against it. Those tests did their job — they confirmed the bridge starts up, accepts data, and produces actions. But each scenario was only about a hundred steps long. A hundred steps is a sprint, not a marathon.
The risk with a translator-style component is drift: a tiny, per-step discrepancy that nobody notices in a sprint but that quietly piles up over a marathon. If the bridge undercounts by one cell every few hundred steps, a short test will never see it. A full training episode — thousands of steps, across several different rooms — will.
So the plan was simple: take our own reference policy, run it through full-length episodes, and at every step compare what the bridge believes about the world against what our own environment believes. Same drone, same path, two independent bookkeepers. If they ever disagree, the bridge has a bug.
2. What the bridge actually does
It helps to picture the bridge’s job as a loop. On each tick it:
- Takes the sensor data coming out of Gazebo.
- Packs that data into the exact observation shape our policy was trained to read.
- Asks the policy what to do, and gets back one move from a small menu (eight possible moves).
- Translates that move into motor commands the simulator understands.
- In parallel, marks cells on its own visited grid as the drone passes over them.
- Computes coverage as: cells visited divided by cells that are open floor.
One of those eight moves deserves special mention. Call it the “fly forward until you hit something” move. When the policy picks it, the bridge tells the drone to keep going straight until a wall stops it — and then the bridge is no longer steering. The drone coasts across many cells on its own. The bridge’s only job during that glide is to watch: it receives position updates from odom roughly thirty times a second and ticks off each cell the drone crosses.
That watching step is exactly where a subtle bug can hide, which brings us to the interesting part.
3. How the check was run
A standalone script imports the bridge’s own grid-and-coverage code directly — not a reimplementation, the actual modules the bridge ships. It then runs that code against a stand-in for Gazebo: a mock environment that behaves like our training world but exposes the same accessors a real Gazebo would expose through ROS2 (Jazzy) topics. Using a mock here means the test is fast and perfectly repeatable, which is what you want when you’re hunting for a one-cell-in-a-thousand discrepancy.
At every step the script compares two things:
- Coverage, as the bridge computes it (from the room’s free-space image), versus coverage as our environment computes it (from its own internal map).
- The visited grid, as the bridge builds it (from the drone’s reported pose), versus the grid our environment maintains internally.
If both the coverage numbers and all 4,096 grid cells agree at every step, the bridge is faithful.
4. The trap: getting to 100% took three tries
This is the genuinely instructive part of the work, and it’s a nice illustration of why full episodes catch what short ones miss.
First attempt. The bridge’s grid was updated only once per environment step. That sounds fine — until the “fly forward until you hit something” move enters the picture. On that move the drone crosses eight to fourteen cells in a single environment step, but the mock environment only reports the final landing position. The bridge saw one update where it should have seen ten, so it marked only the last cell of each glide and ignored everything in between. The two bookkeepers agreed on a dismal 17–23% of cells. Clearly broken.
Second attempt. The fix was to interpolate: between the old position and the new one, fill in the cells the drone must have passed through. Sample a handful of intermediate points and tick those cells too. On straight headings — due north, south, east, west — this worked flawlessly. But on a diagonal heading the naive sampling count came up a couple of cells short, because measuring “how far did it move” by the larger of the horizontal and vertical distances undercounts a diagonal. Agreement jumped to 96–97%, with the bridge undercounting coverage by about one and a half percentage points. Much better, still not exact.
Third attempt. The realization that fixed it: our environment always steps by a fixed distance of one cell-length regardless of direction — a diagonal step and a straight step cover the same true distance. So the correct number of samples to fill in is the straight-line (Euclidean) distance between the old and new positions, divided by the cell size. Use that, and every interpolated cell lands exactly where the environment’s own bookkeeping put it. Agreement reached 100%, bit-exact.
The lesson generalizes nicely: when you reconstruct a path from sampled positions, the number of points you need to fill in is governed by the real distance traveled, not by a shortcut measure that happens to be right only on the axes.
5. Results
With the interpolation done correctly, the full run is clean across the board: nine episodes total — three maps × three random seeds, each three thousand steps long — every one bit-exact.
| Map | Mean env coverage | Spec range @3k | Min match |
|---|---|---|---|
| empty_6x6 | 0.9503 | [0.92, 0.97] | 1.0000 |
| pillar_center | 0.9502 | [0.88, 0.95] | 1.0000 |
| two_chambers | 0.9417 | [0.65, 0.85] | 1.0000 |
The “min match” column is the punchline: across every cell of every episode, the bridge’s grid and the environment’s grid agreed 100% of the time. As a further cross-check, the count of open floor cells computed from the room image matched the count from the internal map exactly on all three maps — 3,844, 3,744, and 3,636 cells respectively.
One honest caveat about these coverage numbers. They are a ceiling, measured under the idealized conditions of the mock environment. A real Gazebo run will land somewhat lower — roughly five to fifteen percentage points — because real physics adds inertia, sliding, and sensor drift that the mock doesn’t model. The expected-coverage ranges in the spec already account for this: the upper bound is the mock ceiling, and the lower bound is the threshold below which we’d want to revisit training.
6. What this means for running in Gazebo
The whole reason to worry about the interpolation logic was the fear that the bridge might miss cells during fast glides. The numbers say it won’t.
In Gazebo the bridge receives the drone’s position through odom at thirty to fifty updates per second. On the “fly forward” move the drone travels at roughly half a meter per second, which works out to entering a new cell about every fifth of a second — five cell crossings per second. With odom running at thirty-plus updates per second, there’s roughly a sixfold safety margin: the bridge gets several position reports for every cell the drone crosses, so it can’t skip one. The bridge is architecturally sound for Gazebo.
There is one practical thing worth keeping in the back of one’s mind. If a recorded session (a rosbag) ever shows the bridge reporting noticeably less coverage than what Aleks can plainly see the drone covering on the Gazebo screen, the first thing to check is the odom update rate for that session — ros2 topic hz /mavros/local_position/odom. A dropped update rate would be the obvious culprit, and that’s a recording problem, not a bridge problem.
7. Takeaway
Two independent components — built by different people, computing the same coverage from different inputs — now produce identical results down to the last cell. That’s the strongest form of confidence you can get short of flying the hardware: not “they look the same,” but “they are the same.” The bridge is cleared for full Gazebo use, with one cheap diagnostic in our pocket if a future recording ever looks off.