claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

32 — Keeping Gazebo Alive: SITL-Only Restarts Without the EGL Death Loop

Restart only SITL between episodes while Gazebo stays alive, sidestepping the EGL-cycle crash that froze the drone at z=0.21.

stablesimulationupdated 2026-06-16T00:00:00.000ZClaudeDroneSimulationDevLog

What this is: The story of why our training runs kept dying after a handful of restarts, and the redesign that let the simulator survive long training sessions by restarting only the flight software instead of the whole world.

Why it’s here: Long autonomous-flight training needs to reset the drone between attempts hundreds of times. If every reset slowly poisons the graphics stack, the run is doomed. This is the engineering log of how we found the poison and routed around it.

Date: 2026-06-09 (night session) Ticket: SITL-RL fine-tune sprint


Glossary

Before the details, here are the moving parts in plain English. If you already live in this stack, skip ahead.

  • Gazebo — the 3D physics simulator. It pretends to be the real world: gravity, collisions, the drone’s body, the ground. When we say “the world,” we mean the scene Gazebo is running.
  • SITL (Software-In-The-Loop) — the real flight-control software running on your computer instead of on a physical flight controller. The drone’s “brain,” simulated. It talks to Gazebo as if Gazebo were reality.
  • mavros — the translator between the flight brain (SITL) and ROS2. It turns flight messages into ROS2 topics that the rest of our code can read (position, attitude, etc.).
  • ROS2 (Jazzy) — the robotics middleware that ties everything together with topics and services. We use the Jazzy release.
  • EGL — the low-level interface between software and the GPU for rendering. Think of it as the handshake your program does with the graphics card before it can draw anything.
  • EGL cycle — what happens every time Gazebo launches: it grabs a fresh GPU handle through EGL. Do this too many times in a row and the driver gets confused and breaks.
  • Liveness check / “gz-alive” — a strategy where we deliberately keep one thing alive and untouched (Gazebo) while restarting other things around it. The opposite of “turn it all off and on again.”
  • hard reset — fully wiping the episode back to a clean start. Historically this meant relaunching everything, including Gazebo.
  • soft reset — a gentler reset: the drone stays armed and airborne, and we just fly it back to the start position instead of landing and taking off again.
  • lockstep — Gazebo and SITL advance time in tight synchronization, one step waiting for the other. Great for accuracy, but it means each side can stall the other.
  • odom (odometry) — the drone’s live estimate of where it is and how it’s moving. Without odom, the control loop is flying blind.

1. The problem: a death loop, one restart at a time

We were trying to run a long fine-tuning session for the flight policy. It kept dying. Not at the start — partway through, reliably, after the simulation had been resetting itself for a while.

The root-cause analysis from the previous session pinned the binding constraint: EGL degradation. Here is the chain of events:

  1. Every hard_reset did a full relaunch — including restarting Gazebo.
  2. Every Gazebo restart triggered a fresh EGL cycle (a new GPU handshake).
  3. After roughly 6 to 8 cycles, the NVIDIA EGL layer broke. The log filled with libEGL: failed to create dri2 screen.
  4. Once that happened, physics stopped stepping. The drone would report itself as armed, but it never actually left the ground — it sat at z ≈ 0.21 m, frozen.

So the simulator wasn’t crashing loudly. It was quietly poisoning its own graphics stack, one reset at a time, until flight became impossible.

The fix direction, agreed with Aleks, was a combination of two ideas: restart only SITL, and keep Gazebo alive. If Gazebo never restarts, the EGL cycle never happens, and the root cause is gone. We layered defense-in-depth on top: a high reset interval, the safety guard turned off, and crash-recovery from earlier work.


2. What we built

2.1 A “keep Gazebo alive” restart mode for SITL

We added a new restart mode that tears down and rebuilds only the SITL side of the stack — the simulated vehicle, the autopilot, and the proxy that connects them — then waits for flight-readiness and brings mavros back. Gazebo is never touched. Its process ID stays the same across the whole operation, which is the whole point: same process, no new EGL handshake, no degradation.

A few details that mattered:

  • World name auto-detection. Instead of guessing the world’s name from the scene file, we ask the live Gazebo what it’s running and address the reset by that name. The name inside the scene file and the file’s own name aren’t always the same, so asking the running instance is the only reliable way.
  • Repositioning the drone. We move the drone back to its start pose through a Gazebo service call rather than reloading any assets. Reloading would mean another expensive (and risky) graphics operation; repositioning is cheap and leaves the GPU handshake untouched.
  • Graceful waiting. The Gazebo-side plugin that bridges to SITL holds its socket open. When SITL momentarily disappears, Gazebo doesn’t crash — it simply waits, physics paused, with no EGL cycle and no error. When fresh SITL reconnects, things resume.

2.2 Reset logic with a fallback and an episode log

The reset function now defaults to the gz-alive path. But it isn’t naive — it falls back to a full relaunch in two situations:

  • Genuine EGL degradation is detected (the dri2 failures cross a threshold), or
  • The SITL restart itself fails (the panel got closed, returns a non-zero code).

We also added a per-episode log written to disk. Each line records the episode number, timestamp, step count, crash count, cumulative reset counts (broken out by SITL-only vs. full relaunch), the worst tilt seen, the lowest altitude seen, and the EGL cycle count. It’s written at the start of the next episode and again on shutdown, so even a run that dies leaves a trail.

2.3 A pre-flight health gate

Before any training run, a health-check script now runs and returns exit code 0 if everything’s fine, or the number of problems found. It checks that the build is clean, Gazebo is alive, mavros is connected, the observation topics are actually publishing, the odometry rate is healthy, the safety guard is off, and EGL is healthy. If the stack isn’t ready, training doesn’t start. No more discovering a broken stack ten minutes into a run.


3. The traps we hit (live, and fixed)

This redesign looked simple on paper. It was not. Four traps, all found the hard way:

  1. Killing a panel doesn’t kill its children. When we killed the foreground process, its child processes — the autopilot and the proxy — survived as orphans and kept holding their network ports. The fresh autopilot then couldn’t bind. Fix: explicitly kill the orphaned children by their command line.

  2. The self-destructing kill. Our wrapper embeds the entire command inside a shell invocation, so the sleeping wrapper’s command line literally contained the text of the launch command. A naive pattern-match kill was matching the brand-new panels and shutting them down instantly. Fix: match only the real binaries, whose exact names don’t appear in the wrapper text. (A subtle save: the binary’s lowercase name differs from the capitalized name in the wrapper, so case-sensitive matching is safe.)

  3. mavros ordering is critical — a double trap.

    • Trap A: If mavros starts at the same time as the SITL restart — before the flight controller is ready — it gets stuck connected: false forever.
    • Trap B: If we instead leave the old mavros running, it reconnects via heartbeat, but it never re-requests the data streams. The result: the position/odometry topic goes dead. (The autopilot version we use ignores the legacy stream-request mechanism; the per-message interval request that actually works is only sent by mavros when it starts up against a ready flight controller.)
    • The solution: restart SITL → wait for flight-readiness (we watch the panel output for a known readiness marker, clearing the old history first so we don’t match a stale one) → then restart mavros. A fresh mavros against a ready flight controller requests full streams, and odometry comes back to life.
  4. Fake sleeps. A “wait” that reads from an empty input source returns instantly — that’s harness pacing, not a real delay. Inside the scripts, we use real sleeps.


4. Acceptance results (live, on the D2 machine)

The headline numbers:

Test What it checked Result
A. gz-alive + EGL-flat 10 consecutive SITL-only resets 10/10 PASS
B. odom recovery + flight One gz-alive reset, then a full takeoff smoke test PASS
Pre-flight gate Health check on a healthy stack exit 0

In test A, the autopilot’s process ID changed on every reset (proving it was a real restart), while Gazebo’s process ID never changed and the EGL failure count never grew — Gazebo never restarted, so the GPU was never re-poisoned. The death loop’s root cause was removed.

In test B, a single gz-alive reset left Gazebo untouched, restored the connection, and brought odometry back to 8–9.5 Hz (versus completely dead when we kept the old mavros running). On the restarted stack the drone went through its full sequence — switch to guided mode, settle the state estimator, arm, take off, climb to z = 1.85 m, hold a stable hover — and passed. The readiness-marker timeout of 90 seconds is acceptable given how rarely a hard reset fires during real training.


5. Two follow-up fixes

5.1 A relative EGL detector

A blocker surfaced before the long run: our degradation detector was counting the absolute number of dri2 failures. But a perfectly healthy Gazebo startup emits roughly ten of these as normal startup noise — already above our threshold. That meant every reset would have fallen back to a full relaunch, dragging us right back into the death loop.

Fix: capture a baseline at construction time, when Gazebo is known healthy, and measure degradation as the delta above that baseline. Now the count reads zero on a healthy stack and after a gz-alive reset, and only rises on genuine degradation. We re-capture the baseline after any full relaunch. (We flagged the same bug to the training side — their own EGL detector needed the same relative treatment.)

5.2 Reposition instead of world-reset

Another blocker: thirteen gz-alive resets in a row all ended connected: false and takeoff became impossible. Root cause: the Gazebo “reset everything” call resets the bridge plugin’s lockstep state at the exact moment no SITL is connected (old one dead, new one still waking up). The fresh SITL then deadlocks — Gazebo frozen waiting for control, the autopilot alive but not heartbeating because it’s waiting for Gazebo.

Diagnosis was clean: skipping the reset entirely produced a clean reconnect, proving the reset was the culprit. The fix (Aleks’s variant) was to reposition the drone with a set-pose service call — using the spawn pose from the scene file — instead of a full reset. Set-pose moves only the model. It doesn’t touch sim-time or lockstep, so there’s no deadlock and the estimator’s origin stays consistent.

Acceptance: 8/8 resets reconnected with live odometry (versus 13/13 failures with the reset call), and the takeoff smoke test passed. A “model-only” variant had been tried earlier and rejected because it tripped a proximity pre-arm check.


6. Postscript: where gz-alive ran out of road, and the pivot to soft-reset

Here’s the honest ending. A final smoke gate (run as recovery after a machine reboot, on Aleks’s direction) exposed a fundamental limit of the gz-alive strategy itself.

  • Run 1 — a baseline full-relaunch flight smoke — passed cleanly (climb to z = 1.85, everything arrived).
  • Run 2 — a combined test cycling through [gz-alive reset → re-takeoff → maneuvers]: cycle 0 (the first reset) passed and climbed to 1.85. But cycles 1 and 2 (the second and third resets in a row) timed out at z = 0.21. Takeoff was accepted, the drone armed, the estimator settled, tilt was zero — but it never left the ground.

Crucially, this was not EGL this time. Gazebo’s process was stable and the failure count was flat. The problem lived deeper, in the flight-control-to-Gazebo path after a repeated SITL reconnect to a still-running Gazebo. In short: gz-alive reposition gives a clean re-takeoff exactly once.

Aleks’s call: don’t keep digging into that low-level synchronization. Pivot to soft-reset instead.

  • Scheduled hard resets: off.
  • Between episodes: soft-reset only — the drone stays armed and airborne and simply flies back to its start position via a position setpoint, no land-and-take-off-again.
  • Hard reset reserved for an actual crash (the drone losing its armed state mid-flight).
  • This works because the state estimator doesn’t meaningfully degrade in our setup without a restart — GPS keeps it corrected.

Validation: a fresh relaunch followed by five episodes (one full takeoff, then four soft-resets) passed 5/5, with a hard-reset count of zero, every maneuver arriving, steady tilt under 0.35°, and no crashes. The long run was unblocked.


7. Takeaways

  • The real fix wasn’t a bigger hammer — it was hitting less. The death loop came from restarting Gazebo too often; the durable answer was to restart it as rarely as possible and, in the end, almost never.
  • Keeping one component alive shifts the bug, it doesn’t always kill it. gz-alive removed the EGL death loop but surfaced a deeper synchronization limit after repeated reconnects. Knowing when to stop digging and route around a layer is its own engineering skill.
  • Health gates pay for themselves. A pre-flight check that fails fast beats discovering a broken stack ten minutes into a multi-hour run.
  • Ordering is everything with the flight stack. Half of our traps were really one lesson: restart mavros after the flight controller is ready, never before, never leave it stale.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR