What this is: the story of building a thin communication adapter that connects our training environment to a real flying drone simulated inside Gazebo SITL, and the hands-on test flight that proved every wire was connected correctly. Why it’s here: training a policy against a fast, simplified grid world is one thing; making that same policy actually take off, rotate, and fly forward inside a physics-accurate simulator is another. This entry documents the bridge between those two worlds and the surprises we hit along the way.
Date: 2026-06-09 Ticket: SITL training sprint, Phase 3 (the final blocker on the simulation side)
Glossary
Before we dive in, a few terms come up constantly. Here they are in plain English with everyday analogies.
| Term | Plain-English meaning |
|---|---|
| SITL | “Software In The Loop.” Instead of a real drone with real motors, the flight controller runs as a program on your computer and thinks it’s flying. It’s a flight simulator for the autopilot brain itself — like a flight-school cockpit simulator, but for the drone’s software. |
| Gazebo | A 3D physics simulator. It gives SITL a virtual world with gravity, walls, and sensors, so the simulated drone actually rises and bumps into things instead of flying through a blank void. |
| MAVLink | The standard “language” drones speak. Commands like arm, take off, go to this point are sentences in this language. Think of it as the air-traffic-control radio protocol between software and aircraft. |
| mavros / MAVROS | The translator that turns MAVLink messages into ROS2 messages and back. Our code speaks ROS2; the autopilot speaks MAVLink; mavros sits in the middle. |
| ROS2 (Jazzy) | The robotics “operating system” everything runs on — a postal service that lets separate programs (sensors, controllers, planners) send each other messages. We use the Jazzy release. |
| SITL comm adapter | The small piece of code this entry is about: a thin connector that lets our training loop drive the simulated drone. It doesn’t reinvent flying — it reuses parts we already had and just exposes them in the shape the trainer expects. |
| Setpoint | A “go here” target you keep sending to the autopilot. The drone continuously steers toward whatever the latest setpoint is — like holding a dog treat where you want the dog to walk. |
| EKF | The autopilot’s internal estimate of where it is and how it’s tilted, blended from many sensors. It needs a few seconds to “settle” before it trusts itself enough to fly. |
1. The problem in one sentence
Our policy was trained in a fast, abstract grid world. To fine-tune it against real physics, it needs to control a drone living inside Gazebo SITL — and nothing yet connected those two. This entry builds that connector.
The training side was already finished by the reinforcement-learning team: the reward logic, the training hooks, the configuration, and the environment class that expects a flying drone to talk to. The only thing missing was the simulation side of the bridge. That was the last blocker of the whole sprint.
2. What we built
The new piece is a class called MavrosSITLComm. It implements a small agreed-upon interface (a SITLComm “protocol”) that the training environment knows how to call. When training starts, this object gets handed to the environment, and from then on the trainer can say “take off,” “do action 4,” “tell me the current state,” and the adapter makes it happen on the simulated drone.
The key design choice: it’s a thin wrapper, not a rewrite. Rather than re-implementing how to read sensors or how to move the drone, it reuses components we already had:
- The observation builder — the same sensor subscriptions we already used (perimeter distance readings, the forward sweep sensor, and odometry).
- The action executor — the existing piece that turns one of the eight discrete actions into a stream of position setpoints at 10 times per second.
- The wall-blocking logic — the same “is there room to move?” calculation the training environment uses, so simulator and trainer agree on what counts as blocked.
Take-off, landing, and arming are driven from inside the adapter using the same constants and sequence our standalone take-off node uses (let the EKF settle for 8 seconds, issue the take-off command, climb to at least ~1.8 meters, stabilize the hover, then hold position). Crucially, the standalone take-off node is not running as a separate process during training — there can be only one owner of the setpoint stream, and that owner is now the adapter.
Threading, briefly
The adapter runs its own ROS2 node on a background daemon thread, where all the sensor callbacks and the 10 Hz maintenance loop live. The training loop calls start_episode, execute, and read_state from the main thread, and those calls block until done. The one place the two threads touch — swapping the current movement target — was already built to be atomic, so there’s no tearing.
What the adapter promises the trainer
This is the contract. The training environment can rely on exactly these behaviors:
| Call | What it returns / does |
|---|---|
read_state() |
Current distance readings (six perimeter values plus the forward sweep), the drone’s pose (x, y, heading), and the commanded servo angle. Distances are raw sensor-frame meters — the environment applies its own corrections. |
execute(action) |
Performs one action and reports back: how many grid cells it traveled, whether it was blocked (tried to move into a wall — we don’t actually fly into the wall; it’s both a safety net and a parity match with training), and whether it crashed (disarmed mid-flight, tilted too far, dropped in altitude, or left the allowed box). |
start_episode(...) |
Resets for a new episode (see the big lesson below). |
hard_reset() |
Relaunches the whole simulation stack, or as a fallback lands, disarms, and warns. |
close() |
Cleanly shuts down the thread, node, and ROS2. |
3. The big lesson: re-takeoff after landing silently fails
The very first two-episode test flight exposed something we did not expect.
Taking off again after a full land-and-disarm cycle does not work in SITL. Episode 1 — a fresh stack — took off perfectly. Episode 2 — which landed, disarmed, re-armed, and tried to take off again — did not climb. The drone reported itself as armed, but its altitude stayed pinned at about 0.21 meters. The take-off command was accepted, but nothing rose.
The fix was to redesign the soft reset so we never land between episodes. As long as the drone is healthily hovering (armed, and above 40% of target altitude), starting a new episode just repositions it in the air — it flies to the new start point and re-stabilizes. A full take-off from the ground only happens on a genuinely fresh stack: the first episode, or right after a hard_reset relaunch. A relaunch gives a clean re-take-off precisely because it clears the autopilot’s accumulated position-estimate bias.
This lined up neatly with a setting we already had: a periodic full stack reboot. That periodic reboot turns out to be the only reliable way to get a clean re-take-off.
There’s one consequence worth flagging for the training side: after an episode where the drone actually crashed, a soft re-take-off is unreliable, so that reset must go through a full hard_reset relaunch. For a normal soft reset where the drone never fell, everything stays clean.
4. The flight-smoke test
Aleks set a simple acceptance rule: a manual flight that does reset → take off → step → hover with no learned model attached. The point isn’t to fly well — it’s to prove the plumbing is connected. The script for this is smoke_flight.py, run against an empty 6×6 room, headless.
Results of the first passing run:
- Episode 1 — take off from the ground: entered guided mode, let the EKF settle for 8 seconds, armed, took off, climbed to 1.85 m, hovered for 5 seconds. The scan action swept the servo from 90° to 120°, a rotate action turned about +14.4°, and a forward action moved — no crash.
- Episode 2 — soft reset in the air: repositioned to a new point, flew there, and held a stable hover.
- Safety guard cooperation confirmed: the deployment-time safety system triggered on a sharp turn during repositioning, and the adapter handled it correctly — held position, re-initialized the target, recovered, and returned to hover.
- Final state: armed, altitude 1.85 m in-band — PASS.
Two known observations (not blockers)
- Arrival tolerance is about one cell wide. The “you’ve arrived” tolerance (0.08 m) is nearly the size of one grid cell (0.1 m), so translation moves often read as zero cells traveled — the drone genuinely arrives within 0.08 m of a target 0.1 m away. This doesn’t affect the reward (translations don’t use that count). It’s an honest physical signal, and fine-tuning is exactly what adapts the policy to this kind of simulation-versus-reality gap.
- Horizontal speed is slow at low altitude. A known characteristic of the navigation controller: at low altitude, forward motion within the time budget covers less than a full cell. That gap is, again, precisely the thing fine-tuning is meant to close.
5. Tests and build
- A unit-test file covers ten pure helper functions: converting a spawn cell to meters (the inverse of the environment’s own conversion), rounding displacement to cells, and computing tilt from orientation. This logic is ROS2-free, so the tests skip cleanly when ROS isn’t installed.
- The full build passed, and the test suite ran 97 passed. (Three errors come from a library that only exists in the training virtual environment, not in a clean ROS setup — not a regression.)
6. First long run died — and what we changed
When the training team launched the first real long fine-tuning run, the process died after roughly 24 minutes. The post-mortem was illuminating.
- The deployment-time safety guard fired constantly — over 1,900 times — aborting rotations and putting the system into a livelock. In the training world, rotations are always valid, so the agent had never learned to deal with a guard vetoing them.
- On one reposition, the safety guard paused the setpoint stream; without a target to hold, the drone sagged, tilted to nearly 62°, dropped, and the autopilot disarmed it.
- The crash-recovery path then couldn’t re-arm, threw an unhandled error, and that error punched all the way up through the training framework and killed the run.
The root cause: the safety guard is a deployment layer that simply does not exist in the training environment. The agent never learned around it, so it behaved unpredictably — aborting rotations (livelock) and pausing the stream (sag). And wall protection during fine-tuning was already consistent between the two sides via the action mask and the blocking logic, so the safety guard was redundant and it broke that consistency.
The fix (Aleks’s chosen option):
- The launch file got a
safety_guardargument (defaulting to on, so deployment is untouched), gated by a condition. Training launches with it turned off. - The launch script got a
--no-safety-guardflag to flip that switch easily. - The adapter’s episode start gained a retry loop: if re-arm or take-off fails, it does a
hard_resetrelaunch (a fresh stack being the only reliable re-take-off and a clean estimate) and retries, only giving up after several attempts. Training no longer dies just because a drone fell.
A fresh flight-smoke run with the safety guard off confirmed the fix: the guard node was absent, three rotations all arrived with zero aborts (livelock gone), a forward step and a longer travel action both completed, the soft reset in the air was clean with no sag, and the run ended armed at 1.81 m — PASS.
7. The binding constraint we hit: graphics degradation
After a second long run died, three more issues surfaced. Two had clean code fixes:
- Boot-gate fix. The deepest cause of the death was stale state: the readiness check saw a “connected = true” flag left over from a killed mavros and passed instantly, racing ahead of a cold boot. The fix resets the cached state before each relaunch and adds a three-level readiness gate (connected → out of the initializing mode → odometry reporting a valid altitude) with a minimum 35-second wait. Validation confirmed the stack now reports ready only when it genuinely is.
- Early-crash detection. A background latch now flags a crash (extreme tilt, or airborne but near the floor) before the autopilot disarms, so a crash inside a long action is caught immediately. Validation confirmed no false crashes in normal flight, with correct triggering on the ground.
The third issue is the real wall, and it lives outside our code entirely. Each full hard_reset relaunches Gazebo, and relaunching Gazebo repeatedly degrades its hardware graphics layer. After roughly six relaunch cycles, Gazebo’s rendering breaks (a driver failure in the graphics back end), and a broken Gazebo stops stepping the physics — so the drone reports itself as armed but never rises off the floor. Since every hard_reset is one of those cycles, any relaunch-based recovery will inevitably hit this ceiling after about six resets, no matter how good the code is.
The strategy going forward, pending sign-off:
- Priority: a
hard_resetthat doesn’t kill Gazebo. Restart only the autopilot SITL process (and mavros), and leave Gazebo running. The position-estimate bias lives in the autopilot, not in Gazebo, so restarting just the autopilot clears the estimate without a graphics cycle. - Minimize resets overall — fewer crashes (thanks to the guard being off and the new crash recovery) means fewer relaunches.
- A clean reboot of the test machine for valid re-take-off validation (Aleks’s area).
8. Where this leaves us
The simulation side of the bridge is built, tested, and proven by hand. A learned policy can now take off, rotate, scan, and fly forward inside Gazebo SITL through a thin adapter that reuses our existing flying components and presents exactly the contract the trainer expects. The remaining open work is on the recovery path — making resets cheap enough that long training runs don’t grind against the graphics-cycle ceiling.