claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

21 — v2 Block 1: Auditing Links, Nodes, and Protocols

A full static and runtime audit of the ClaudeDrone simulation stack — node links, QoS, topic rates — uncovering four silent bugs and fixing them.

stablesimulationupdated 2026-06-16T00:00:00.000ZClaudeDroneSimulationDevLog

What this is: The first block of a “v2” cleanup effort, where I stop tuning behavior and instead audit the plumbing — every connection, message, and timing rate in the simulation stack — to find out what was quietly broken underneath.

Why it’s here: Before you can trust what an autonomous drone does, you have to trust what it sees and hears. This entry documents how four silent wiring bugs were starving the system of correct data — and how each one was tracked down and fixed.

Date: 2026-06-06 Ticket: v2 Block 1 — links & protocol audit (branch v2_takeoff_plus_model)


Glossary

A few plain-English terms before we dive in:

  • Links audit. Think of the drone’s software as a city of little workers (nodes), each shouting messages onto named channels and listening on others. A links audit is walking the whole city and checking that every shouter is heard by the right listener — that nobody is shouting into an empty room, and nobody is straining to hear a channel that doesn’t exist.
  • Protocol. The agreed-upon “rules of the conversation” between two nodes: not just what channel they talk on, but how — how reliably messages are delivered, how often, and in what format. Two nodes can be on the same channel and still fail to talk if their protocols disagree.
  • Message flow. The path a single piece of information takes from where it’s born (a sensor reading) to where it’s used (a decision). If any link in that chain is broken or stale, the decision at the end is made on bad information — even though nothing visibly crashes.
  • QoS (Quality of Service). ROS2’s settings for how a channel behaves — for example “RELIABLE” (resend until it arrives, like a registered letter) versus “BEST_EFFORT” (send once, don’t worry if it’s dropped, like a live radio broadcast). A sensor that floods data wants BEST_EFFORT; a one-time “ready” signal wants RELIABLE. If a publisher and subscriber pick incompatible settings, they silently never connect.
  • Topic rate (Hz). How many times per second a channel carries a message. 10 Hz means ten updates a second. Too slow and the drone reacts to stale data; far too fast and you waste CPU re-sending things nobody needs that often.
  • DDS. The low-level networking layer underneath ROS2 that actually moves messages between nodes and enforces those QoS rules. It’s the postal service the whole city runs on.

1. What we wanted

The previous sprint — getting a learned model to drive the drone in simulation — had stalled. Rather than keep tuning behavior on top of a foundation I didn’t fully trust, Aleks proposed a systematic v2 plan: first verify every connection and protocol, then map out exactly how data reaches the decision-making model, then build clean isolation layers between the pieces.

This entry is the first block of that plan: the audit. The goal was simple to state and tedious to do — go through the entire stack and confirm that what should be connected actually is connected, at the right rate, with the right protocol.

2. What we did

I ran the audit in two passes.

Static pass. I read through every node in the two main subsystems — the drone simulation nodes and the bridge that connects the model to the simulator — and listed, for each one, what channels it subscribes to, what it publishes, its QoS settings, its rates, and its timers. On paper, this is where naming mismatches and protocol disagreements jump out.

Runtime smoke test. Reading code only gets you so far; some bugs only show up when everything is actually running. So I brought up the full stack headless (no graphics window) on machine D2 — full launch with mavros, an empty 6×6 indoor room world, indoor flight parameters, fresh EEPROM. Then I inspected the live system with ROS2’s introspection tools: one to dump each topic’s publishers, subscribers, and QoS, and another to measure the actual message rate on the wire.

The combination matters. The static pass tells you what should happen; the runtime pass tells you what is happening. The interesting bugs live in the gap between the two.

3. Results — four silent bugs, all fixed

Every one of these had the same nasty property: nothing crashed, no error appeared, the drone kept flying. The data was just quietly wrong.

Bug 1 — A safety parameter that was being ignored

The launch file was handing the safety node a parameter under one name, but the node was actually declaring a different name for that setting. Because the names didn’t match, the value was silently dropped on the floor and the node fell back to its built-in default. So the operator thought they were tuning the safety stop distance from launch, but in reality that knob was dead — the node had been using its default the whole time.

The saving grace: an earlier tuning effort had (unknowingly) been calibrated against that real default, so the actual flight behavior never changed. But the launch file was lying, and the tuning path through launch was broken.

Fix: use the correct parameter name and set the value explicitly. Verified with a live parameter query that the node now reports the intended value.

Bug 2 — The model never knew where its sensor was pointing

This was the most consequential one. One of the drone’s distance sensors sits on a small servo that sweeps back and forth. The observation builder — the node that assembles what the model “sees” — needs the servo’s angle to know which direction that distance reading is pointing. Without it, you have a range but no bearing: “something is 1.2 meters away” with no idea of where.

The builder was listening for the servo angle on a default channel name, but the bridge that translates Gazebo’s data into ROS2 publishes it on a much longer, world-specific channel name. The override that would have pointed the builder at the right channel was never actually being passed through. Result: the model never knew the servo angle — that input was a constant zero. A strong suspect for the erratic behavior we’d been chasing.

Fix: add a proper launch argument for the servo-angle channel, defaulting it from the active world name (the same mechanism already used for the world itself), and pass it down into the node. At runtime the channel now exists and publishes correctly.

Bug 3 — Position updates throttled down to a trickle

The drone’s local-position odometry — how the navigation logic knows where it is — was arriving at only about 3.8 updates per second, because the SITL telemetry proxy ships at a low default stream rate. At the drone’s cruising speed, that’s roughly 8 cm of travel between position updates, while the “you’ve arrived” tolerance was only 8 cm. So the drone would routinely sail past its target between updates and report a false “arrival timeout — distance still 0.10 m, over the 0.08 m tolerance.” That maddening symptom from the previous sprint suddenly had a cause.

Fix: raise the telemetry stream rate in the SITL launch command. Verified: position updates climbed to about 8.8 Hz, comfortably faster than the navigation loop needs.

Bug 4 — A relative path quietly killing SITL

The SITL launch step changes into the ArduPilot directory before starting, but it was being handed a relative path to its parameter file. After the directory change, that relative path no longer pointed at anything, so the flight simulator failed to start — and, worst of all, the rest of the stack came up fine without it, so the failure was easy to miss.

Fix: resolve the parameter path to an absolute path before changing directories. Verified by re-running with a relative path and confirming SITL now starts correctly.

4. What we checked and found healthy

Not everything was broken — and confirming the good parts is just as valuable as finding the bad ones.

  • QoS / protocols. No mismatches. The Gazebo bridge publishes RELIABLE, and the RELIABLE subscriptions on the other end are valid. The mavros sensor topics use BEST_EFFORT, and the nodes consuming them subscribe with the matching sensor-data profile. Lessons from earlier QoS troubleshooting held up.
  • Channel names. The whole chain — from Gazebo, through the bridge, through the sensor monitor, into the observation builder and safety node — lines up correctly (with the single exception of Bug 2 above).
  • Nodes. Every node we expected to be alive was alive: the Gazebo bridge, sensor monitor, servo command, sweep, autoscan, safety node, and the full mavros set. The one-time “takeoff ready” signal, which needs to be available to late-joining listeners, was using the right latching protocol and worked correctly.

Measured topic rates (runtime, headless)

Topic Hz Note
Distance sensor channels 0–5 10.0 normal
Sweep scan 9.8 normal
Perimeter aggregate 58.6 re-publishing the whole array on every channel (6 × 10 Hz) — flagged for Block 2
Local-position odometry 3.8 → 8.8 Bug 3 fixed
Servo joint state 966 no throttle, heavy Python-callback churn — flagged for Block 2

The two bold rows aren’t bugs exactly — the data is correct — but they’re wastefully fast, and I’ve parked them for the next block.

5. Sources

  • ClaudeDrone simulation engineering journal, this entry’s predecessor sprint (the model-to-sim effort that stalled).
  • Runtime introspection captured on machine D2 with the full headless stack: per-topic publisher/subscriber/QoS dumps and live rate measurements.
  • Earlier QoS troubleshooting notes, whose conclusions this audit re-confirmed.

6. What’s next

A couple of threads carry over into the next block:

  • The over-fast channels. The perimeter aggregate at ~58 Hz and the servo joint state at ~966 Hz are far more traffic than anyone needs. I want to solve these together with a bigger design question: what, and when, should the model actually see? Right now the model takes a snapshot of the world on a free-running 10 Hz timer, even though its actions can block for anywhere from half a second to fifteen seconds. The cleaner approach is to snapshot the observation precisely at the boundary of each action, after letting the world settle — rather than grabbing whatever happens to be on the wire at an arbitrary tick.

  • A logged-but-not-yet-fixed conflict. The servo command channel currently has three potential writers: the servo command node, the optional sweep storage node, and the action executor. If autoscan triggers a sweep while an action is also driving the servo, the servo receives conflicting commands. I’ve recorded this but deliberately left it alone — the right fix belongs to a later block, where we’ll also settle who owns the servo.

Block 1 did its job: the foundation is now honest. Four silent bugs that were feeding the model bad data are gone, the healthy parts are confirmed healthy, and the wasteful traffic is documented and queued. With trustworthy plumbing underneath, the harder questions about what the model should perceive can finally be tackled on solid ground.

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR