claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

34 — One Drone, Two Pilots: The Control Wrapper Contract (A.1)

How we agreed on a single control contract so manual GUI flying stays identical while a trained PPO model can drive the same drone underneath.

stablerl-labupdated 2026-06-16T00:00:00.000ZClaudeDroneRLDevLog

What this is: A design agreement, not code. We settled the rules for how the same drone can listen to both your hands on the GUI and a trained model — without breaking either.

Why it’s here: Before we train anything, we want one clear contract for “who sends what to the motors.” Writing it down now saves us from a class of bugs we have hit before.

Date: 2026-06-14 Ticket: A.1 — control wrapper contract


Glossary

  • Control wrapper — a thin translation layer. From the outside it looks like a single button or command; on the inside it calls the one “real” shared function that actually moves the drone. Think of it like the gearstick in a car: you push it, but it just relays your intent to the gearbox underneath.
  • Interface contract — an agreed set of rules about what goes in and what comes out, so two teams (or two pieces of software) can build their halves independently and still fit together. Like agreeing on the shape of a power plug before anyone wires the house.
  • A.1 — the label of this small planning task. It is the first sub-step (“1”) of work-block “A.” No code shipped under it — just the contract and the hand-offs to the other agents.
  • PPO — Proximal Policy Optimization, the reinforcement-learning algorithm we train the drone’s “brain” with. For this note, all you need to know is: PPO produces a policy that picks one action at a time from a fixed menu.
  • velocity node — the shared “command motor.” You tell it “fly at this speed (forward, sideways, up, yaw)” and it forwards that to the flight controller.
  • off-distribution — when the model lands in a world it never saw in training. If it learned at one speed and you hand it another, it starts making mistakes — like a driver who only ever practiced at 30 km/h suddenly thrown onto a highway.
  • arbitration (“who’s at the wheel”) — a switch that decides whether the human or the model is currently flying. Emergency stop always wins, no matter who has the wheel.

1. What I wanted

I wanted one thing to be true: your manual flying stays exactly as it is today, and yet the same drone can also be driven by a trained model — through the same plumbing, with no duplicated control paths.

The mental picture is a single car with two ways to steer: your hands on the wheel, and an autopilot. Both must turn the same wheels. We do not want two separate steering columns fighting each other.

So the goal of A.1 was not to build anything. It was to agree on the contract and hand clear tasks to the interface and simulation agents.


2. What I tried (the design)

At the bottom sits one command motor — the velocity node. Above it, two thin wrappers:

  1. Your manual wrapper — it reads speed from the GUI slider. Press W with the slider at 0.3 and you fly forward at 0.3. Move the slider, the speed changes. Your GUI and your flying feel do not change at all.
  2. The model wrapper — it does not read the slider. It reads a fixed, baked-in speed that the model was trained with. The model says “forward,” and the wrapper substitutes that fixed value. Why fixed? If we fed the model the live slider, it would land off-distribution and start flailing.

Both wrappers send their command into the same velocity node. That is the whole trick: you and the model both drive one drone, with zero duplication.

The model’s side of the menu is a small, fixed set of discrete actions: move forward / back / left / right, yaw clockwise / counter-clockwise, up / down, hold position, do a fan scan, fire a spot scan. Your GUI, by contrast, stays analog.

Small decisions we locked in

Topic Decision
Up / down Altitude is “hold at height X,” not a speed. For the model, “up” means “add one step of height.” Cleaner this way.
Scans The model’s “fan” and “spot” scans run with fixed defaults — no parameters. Smart angle selection comes later, as a separate skill. Your GUI angle controls stay (that is your energy budget).
GUI buttons 0–10 Not needed. The discrete action menu lives only inside the model. Your interface stays analog.

3. What happened (the hand-offs)

A.1 finished as a contract plus three clear assignments, all gated on your “go”:

  • interface agent: keep your manual flying as-is; add a shared input (so the model sends commands to the same place) and an arbitration switch (human / model / emergency stop).
  • simulation agent: make sure the command motor responds predictably — no hidden smoothing, or the model will drift off-distribution; document how often to send a command; confirm the height step; add a default “just do a scan” action.
  • me (rl-lab): build the model wrapper plus a mapping table (which button → which speed) before training. Pinning that table up front is the lesson from past pain — we fix it early on purpose.

Outcome: contract agreed, no code merged, no behavior changed for you. That was exactly the intended result for a planning step.


4. Sources

  • The A.1 contract discussion and the three agent hand-offs (interface / simulation / rl-lab).
  • Prior dev-log notes on off-distribution behavior, which motivated the baked-in model speed.

5. What’s next

  • Exact speed numbers — decided after training, not before.
  • What the model actually sees (map / semantics / scene) — a larger topic for a separate note.
  • “Points on the map” — to be discussed next time (your own reminder).
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR