claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

39 — Locking the Discrete Control Contract for the RL Drone (Scenes S1-S3)

How we froze the shared-velocity, altitude-step, and scan-trigger interfaces so an RL agent can fly the simulated drone safely.

stablesimulationupdated 2026-06-16T00:00:00.000ZClaudeDroneSimulationDevLog

What this is: A walk-through of how we pinned down the exact “control contract” between the reinforcement-learning brain and the simulated drone — the small set of commands the agent is allowed to send, and exactly how the flight stack reacts to each one.

Why it’s here: Before you let a learning agent fly a drone, both sides have to agree on the rules of the conversation. This entry documents the three first scenes of that agreement (we call them S1, S2, S3), and the field tests in Gazebo that confirmed the wiring actually behaves the way the documents promise.

Date: 2026-06-14 Ticket: Track-1 — discrete control wrapper, scenes S1-S3


Glossary

  • Discrete wrapper — a thin software layer that turns the drone’s continuous controls (any speed, any angle) into a short menu of fixed choices, like a TV remote with a handful of buttons instead of a free-spinning dial. The learning agent picks one button at a time; the wrapper translates that button press into a real flight command. We chose this on purpose: a small menu is far easier for an agent to learn than an infinite range of values.
  • Interface contract — a written agreement that says “if you send this message, the drone does exactly that.” Think of it like the menu and prices at a restaurant: both the customer and the kitchen rely on it being stable. If the contract drifts, the agent learns one thing and the drone does another.
  • Shared velocity node — the single piece of software that accepts “fly this direction at this speed” requests. It is “shared” because it does not care who asks — a human with a joystick, a keyboard teleop tool, or the RL agent all use the same door. That is what makes a human-trained skill transfer cleanly to the agent.
  • Watchdog — a safety timer. If the controller stops hearing fresh commands for too long, it assumes something broke and parks the drone in a stable hover instead of coasting blindly.
  • SITL (Software In The Loop) — running the real autopilot firmware as a program on a PC, with no physical drone. The flight brain is genuine; only the body is simulated.
  • mavros — the bridge that lets ROS2 software talk to the autopilot in its own language.
  • Scenes S1-S3 — three slices of the contract we locked in this session: S1 = how the agent commands motion, S2 = how it changes altitude, S3 = how it triggers a sensor scan.

1. The job was to confirm, not to build

The headline finding came early: the infrastructure for all three scenes already existed in the flight stack. So this session was an audit, not a construction project. I read the actual running code — the action executor, the manual-flight node, the sweep node, and the servo command node — and wrote down, scene by scene, the real behaviour into our interface-contract document.

This distinction matters. It is tempting to “improve” things while documenting them. Instead, the rule was: describe what the code truly does today, find the sharp edges, and only change anything later, deliberately, on an explicit go-ahead.

2. Scene S1 — one shared door for motion

The agent sends body-frame velocity requests (forward/back, left/right, and yaw rate) through a single message channel. Whoever publishes to that channel — a human GUI, a teleop tool, or the RL agent — is interchangeable above the controller. Picking which source is allowed to drive at any moment is a separate arbitration concern handled outside the simulator.

A few properties we confirmed and wrote down:

  • Speed limits exist. Horizontal speed and yaw rate are clamped to safe ceilings, so the agent’s chosen amplitudes simply have to live inside those bounds.
  • Vertical motion is automatic. The agent does not steer altitude through the velocity channel; height is held by a dedicated controller. That keeps the motion menu simple.
  • The live path is clean. Once the drone is airborne in manual-flight mode, the held velocity vector is forwarded as-is to the autopilot, which performs a single body-to-world rotation internally. The lesson from earlier debugging was sharp: never rotate by hand on top of that, or you get a double rotation.

The two-path fork — the key catch

The most important discovery was that the code contained two ways to publish velocity:

Path Behaviour Verdict
Clean manual-flight path Forwards the agent’s vector directly, one rotation, no hidden safety re-clamp The one RL must use
Legacy velocity path Added an extra hidden clamp and rotated by hand (risking a double rotation) Effectively dead on the live stack

On the live stack the legacy path was never reached — manual-flight always enters the clean path first. But a dormant second path is a trap waiting for a future change to wake it up, so the recommendation was to remove or gate it as a separate, deliberate commit. (We did exactly that later in the session — see section 6.)

The watchdog (S1.c)

The executor streams setpoints to the autopilot at a steady rate on its own; the agent is not required to stream. But there is a freshness timeout: if no new velocity command arrives within half a second, the controller fails safe into a stable position-hold hover. The RL agent’s “hold this action” duration comfortably fits inside that window, so normal operation never trips the watchdog. To hold a direction longer than the timeout, the agent simply re-publishes before it expires.

3. Scene S2 — altitude as a small step (confirmed)

Altitude changes go through a dedicated target channel: send a desired height in metres, the controller clamps it to a safe band, holds horizontal position, drives vertically to the target, and hands control back once it arrives within a small tolerance. The agent reads back the current altitude target (echoed periodically) and asks for “current ± a step.” How big that step is, is the learning team’s choice — the simulator does not freeze it. Vertical velocity sent through the motion channel is intentionally ignored here, which keeps the two altitude paths from fighting.

4. Scene S3 — triggering a scan with no parameters (confirmed)

There are two scan flavours, both designed so the agent can fire them with the simplest possible signal:

  • Fan scan — a single empty trigger sweeps the rangefinder across a wide arc and reports the result. One catch: the default full-resolution sweep takes a fair while per pass, which is long for a training episode. That became the motivation for a faster variant (section 6).
  • Precise scan — the agent points the servo straight ahead and reads the single forward distance. There was no single atomic “aim then grab one reading” node, so the recommended approach is for the agent to read the existing scan stream after the servo settles, rather than adding a new node.

5. Designing C1 — turning scans into labelled points

After the three scenes, I designed the “C1” contract: how the simulator can turn live scans into structured scan-point records for the learning team’s data pipeline. The geometry is straightforward once you fix the conventions:

  • A scan point is built from the drone’s pose plus the servo bearing plus the measured range. The servo’s zero angle maps to “straight ahead,” and the usable arc covers the front half-sphere (left-through-front-through-right).
  • Each point can be referenced two ways: against the autopilot’s own odometry estimate (which can drift), and against the simulator’s ground-truth pose. We keep both, because comparing them is the whole point of the layer — it is what makes the data useful for noisy, real-hardware cases later.

Then I implemented the scan-points node. A subtle problem surfaced: the ground-truth pose feed is an unnamed list of poses, so I had to identify which entry is the drone. The first guess — nearest by horizontal position — turned out to be ambiguous (more on the fix below).

6. Field verification in Gazebo, and three follow-up changes

I brought up a test stand in Gazebo (its own isolated instance, kept separate from the interface stack) and ran real scans.

The stand fix. The ground-truth pose feed listed dozens of unnamed entries — model roots and every sub-link. The drone has several links sitting near the same horizontal origin, so “nearest by horizontal position” was a coin-flip. The fix was to match in full 3D including height: the drone’s root sits at a distinct altitude that the sub-links do not share. After that, the drone is identified reliably both on the ground and in the air.

Scan-point results. Two sweeps produced clean, sensible geometry:

Sweep Condition Points Notes
s0 on the ground 181 Hit the south and east walls at expected ranges, plus a near interior wall; ground-truth located
s1 hovering 181 Ground-truth vs odometry agreed within about a centimetre; point heights matched the mount offset exactly

The bearing sign and zero were correct (nose pointing east), points landed on the real walls, the sensor mount height checked out, and the front half-sphere coverage was exactly where it should be. The dual pose (odometry plus ground-truth) was captured correctly.

Visual track lives in $DRONE_MEDIA_ROOT/sim/c1-stand-verify-2026-06-14/:

  • c1_points_topdown.png — scan points overlaid on the stand geometry (interior room plus outer walls)
  • sim_20260614_150837.mp4 — the Gazebo flight and sweep

Three concrete changes followed on the same go-ahead:

  1. Removed the dead legacy velocity path. With its only real caller always entering the clean path, the legacy method, its dispatch branch, its private clamp helper, and the now-unused clamp constants all came out. One deterministic velocity path remains. The bridge test suite stayed green (the only errors were pre-existing and unrelated to this change).
  2. Added a fast fan-scan trigger. A coarser sweep using a bigger angular step and shorter settle finishes in roughly a third of the time of the full-resolution sweep (about 6 seconds wall-clock versus the slow default), producing fewer but still well-placed points. It is a separate trigger, so the GUI’s fine sweep is untouched, and it stays tunable. Stand verification confirmed the fan rotates correctly with heading.
  3. Investigated dynamic drift — see the honest negative result below.

Visuals for the fast sweep: c1_fast_sweep_s3.png (fan at two headings, rotating with course) and sim_20260614_152627.mp4.

7. An honest negative result: the simulator is too clean

I tried hard to make the odometry estimate visibly diverge from ground truth, because that gap is exactly what the dual-pose layer is meant to capture. Every attempt failed for the same reason — the simulated autopilot’s state estimator is nearly perfect:

  • Injecting a GPS glitch in hover: the estimator gated it out (rejected the suspicious update); the gap stayed about a centimetre.
  • Injecting GPS noise: filtered away; gap still about a centimetre.
  • Disabling GPS in a stationary hover: zero drift, because dead-reckoning a motionless hover accumulates no error.

The conclusion is that the default SITL simulation is simply too pristine for natural drift to appear. Producing a realistic gap would require either injecting sensor noise plus actual movement, or a full reconfiguration onto an optical-flow style estimator that relies only on relative pose — both well beyond a quick demo. The important part: the dual-pose plumbing is proven correct, and its value is precisely for real-hardware and noisy cases. No numbers were faked.

8. A teardown footgun worth remembering

While cleaning up, the stack-kill script left orphaned simulator and node processes behind, which is a real risk for the GPU. Two root causes: my own teardown commands were accidentally matching and killing their own shell (the search pattern appeared in the command’s own arguments), and the kill script’s patterns did not cover the autopilot, the mavros node, or the bridge nodes. The fix added the missing patterns and an optional session-kill flag, and documented the self-kill trap. Verified in anger: a single invocation tore down seventeen processes and the session in about a second with no self-kill, leaving the GPU clean — something the old version could not do.

9. Where this leaves us

The Track-1 contract (S1-S3) is locked and verified on the stand, the single clean velocity path is the only one left, the fast sweep is in place, and the C1 scan-point node works with the 3D identification fix. The stack-kill script is fixed and battle-tested. Realistic odometry drift remains an open, honest negative for now. Still open with the learning team: the action-amplitude parity table, the scan-precise return format, and the data sink details — all to be settled in the next coordination round.

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR