claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

22 — V2 Block 2: Shaping the Data Flow Into the Policy

How we matched the drone's live sensor stream to what the trained policy expects — settle delays, freshness gates, and quieter sensor rates.

stablesimulationupdated 2026-06-16T00:00:00.000ZClaudeDroneSimulationDevLog

What this is: A maintenance-and-tuning pass on the bridge that feeds our flight policy. We took a hard look at what the model expects to see versus what our sensors actually deliver, and closed the gaps between them.

Why it’s here: A policy trained in a clean 2D world will quietly misbehave on a real drone if the data it reads at decision time doesn’t match its training diet. This entry documents the mismatches we found and the small, surgical fixes that brought the two into alignment.

Date: 2026-06-06 Ticket: V2 Block 2 — model data flow / training parity


Glossary

  • Policy / model. The trained “brain” that, given a snapshot of the drone’s situation, decides the next move. Think of it as a chess engine that only plays well when it sees the board exactly the way it was taught.
  • Observation (obs). The snapshot we hand the policy on each turn — position, coverage map, servo angle, and so on. The single most important thing about an obs is that it must be honest and settled, not blurry.
  • Inference loop. The repeating cycle of “read sensors → build obs → ask the policy → execute the move → wait → repeat.” Our loop ticks at a steady background rate but only decides at the right moments.
  • Turn-based policy. A model trained so that each decision happens only after the previous move fully finishes and everything has stopped moving — like a turn-based board game, not a live video game. This matters a lot for when we’re allowed to read the obs.
  • Blocking control. The matching deployment style: ask the policy once, carry out the move completely, then ask again. The opposite (deciding while still moving) would require retraining the model from scratch.
  • Freshness gate. A guard that refuses to make a decision until the sensor values are recent enough — so the policy never acts on a stale reading from before the drone stopped.
  • Settle. A short, deliberate pause after the drone arrives at a target, letting it stop wobbling before we take the snapshot.
  • Coverage. How much of the room the drone has visited — our main “is it actually exploring?” metric, scored from 0 to 1.
  • Servo / scan servo. A small motor that aims a sensor. Action 6 of the policy controls it exclusively, so nothing else is allowed to move it behind the model’s back.
  • mavros / SITL / Gazebo / ROS2 (Jazzy). Our standard stack: SITL is the software-in-the-loop flight controller, Gazebo is the physics world, ROS2 (Jazzy) is the messaging layer, and mavros bridges flight commands.

1. What we wanted

Aleks framed the block as a plain question: “Think through what the model is waiting for versus what our sensors actually give it. Maybe we’re sending the model data too often. Find the optimal data flow.”

So this wasn’t a feature block — it was a parity audit. My goal was to make the live drone hand the policy the same kind of snapshot, at the same moment in the move, that it learned from during training. Anything else and the model is essentially reading a foreign language.

2. What the model is actually waiting for

I cross-checked our trained policy against the training environment and against the published literature on deploying turn-based policies on real robots.

The verdict was clear and consistent: our policy was trained turn-based. The observation is built strictly after a move completes, when the world is standing still. The right way to deploy such a policy is blocking — decide once at the boundary of each move, and use the fast background tick (about 10 Hz) only to keep sensor buffers warm, not to keep re-deciding.

So “are we sending data too often?” turned out to be the wrong worry. The decision rate was already effectively one-per-move because of how the action group is structured. The real problem was the quality of the snapshot at the move boundary — we were reading it too early, before the drone had truly settled, and some of our sensor feeds were noisier and chattier than they needed to be.

3. What we did — the gaps and their fixes

Here are the mismatches I found between live behavior and training, and how I closed each one.

# The mismatch The fix
1 Intermediate cells along a flight path weren’t being marked as “visited” — the bridge only recorded the final pose once per step, so the coverage map had a gap running along the drone’s own trajectory. The visited-cells update now fires on every arrival-poll, so the whole path gets credited, not just the endpoint.
2 There was no settle pause. The observation was being taken 0–100 ms after arrival — right in the middle of the flight controller’s overshoot wobble. Added a short ~0.3 s settle after every arrival (both moves and turns), which in practice lands around 200–500 ms of genuine stillness.
3 Caches could hand back values from before the drone stopped. Added a freshness gate: a decision is skipped until the relevant sensor reads are confirmed recent; the tick simply re-checks a fraction of a second later.
4 The scan servo started at one angle while the training world assumed another, so the very first observation disagreed with reality. The executor now initializes the servo to the trained starting angle and publishes that command up front.
5 The servo angle reported to the policy came from a high-rate hardware joint feed, but training assumed commanded equals actual (no servo dynamics) — and the raw feed fired hundreds of times per second, churning the inference process needlessly. Switched the reported angle to the executor’s commanded value, matching training and quieting the loop.
6 A background auto-scan routine periodically swept the servo on its own, while the policy still believed the servo was wherever it had last put it. Added a launch flag to disable auto-scan; RL runs are now required to use it, since action 6 owns the servo exclusively.
7 A perimeter sensor was republishing its entire array out of every channel callback, pushing it to roughly 58 Hz with mostly stale entries. Replaced that with a clean 10 Hz timer that publishes once per cycle.

The throughline across all seven: give the policy a calm, honest, settled snapshot at the right moment — and stop flooding the inference process with data it doesn’t use.

4. What we deliberately left alone

It’s worth recording the non-changes too, because “we thought about it and chose not to” is itself a decision:

  • The 10 Hz background tick stayed. It’s cheap, decisions are already gated by freshness and the action structure, so there was no reason to slow it down.
  • No move to continuous / concurrent control. Deciding while still moving would mean retraining the policy with delay-aware dynamics. Our blocking mode is the correct deployment for a turn-based policy, so we kept it.
  • Wall-following phase keeps its own visited-grid mechanism, handled by its dedicated phase controller — no need to touch it.

5. Results

Verification came in two layers.

Build & tests. A clean full build, plus 30 system tests passing and 36 policy-side tests passing — including an integration test that mocks the full RL stack.

Live end-to-end run (headless, empty 6×6 room, policy-only, ~200 seconds):

  • 102 steps, zero errors, zero freshness-skips.
  • Only 3 arrival timeouts total — previously these were near-constant.
  • Coverage climbed monotonically, from 0.005 up to 0.027 — slow but genuinely and steadily increasing.
  • The servo initializing to the trained start angle was confirmed in the log.
  • The drone actually toured the room (finishing at roughly (-2.4, -1.6, z=2.75)) — no flyaway, no getting stuck against a wall.

For contrast, an earlier run before these fixes had stalled with a yaw-drift “dance” pinned against one wall. After the parity pass, that pathology was gone and exploration resumed.

One loose end I also tidied: an auto-resolve path for the free-space map that had previously been a stub, which had been under-counting coverage and making the target unreachable.

6. Sources

These came from focused research passes; full links live in the underlying reports.

  • Thinking While Moving (ICLR 2020) — the blocking-vs-concurrent control distinction.
  • A sim-to-real coverage-path-planning study (arXiv 2406.04920) — real-world observation latency on the order of ~450 ms.
  • robo-gym — the latest-value cache pattern for asynchronous sensors.
  • ROS2 message_filters / topic_tools — considered and rejected as overkill for our observation needs.

7. What’s next

With the data flow now honest and settled, the natural follow-on is to watch coverage behavior over longer runs and richer rooms, and to keep an eye on whether the settle window needs tuning per move type. The parity foundation is in place; the next work builds on top of a snapshot the policy can finally trust.

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR