claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

44 — Flight Dynamics: Tilt, Inertia, and a Conscious Scanner

The drone learns to feel inertia, tilt forward to fly, wobble like a real quadcopter, and aim its TF-Luna fan-scan on purpose.

stablerl-labupdated 2026-06-16T00:00:00.000ZClaudeDroneRLDevLog

What this is: An engineering-journal entry about teaching the simulated drone to move like a real quadcopter — feeling inertia, tilting to accelerate, wobbling before it settles — and letting the agent steer its long-range scanner on purpose instead of letting it spin automatically.

Why it’s here: Up to now the drone flew “perfectly”: ask for a speed and it appeared instantly, no lean, no overshoot. That is not how flying machines behave, and training against a fantasy makes the learned policy useless on real hardware. This entry is where the physics got honest.

Date: 2026-06-15 Ticket: rl-lab dev-log 44


Glossary

  • Inertia — a body cannot speed up or stop instantly. Think of a car: you press the gas and it gathers speed over a second or two; you brake and it keeps rolling for a bit. The drone now behaves the same way instead of teleporting from rest to full speed.
  • Pitch / roll — the drone leaning forward-and-back (pitch) or side-to-side (roll). A quadcopter has no wheels and no rudder; the only way it moves horizontally is to tilt. Lean forward and the rotors push you forward. So tilt is not a side effect of flying — it is flying.
  • Overshoot / wobble — when the drone tilts to a target angle, it sails past that angle and rocks back and forth a few times before settling, like a diving board that keeps bouncing after you step off. Real airframes do this; perfectly damped ones do not exist.
  • Fan-scan / TF-Luna — a long-range single-beam distance sensor mounted on a small servo motor. By sweeping the servo, the drone “paints” the space ahead in a fan shape, the way you sweep a flashlight across a dark room to see what’s there.
  • Conscious scan — an action the agent chooses (where and when to point the beam), as opposed to something the simulator does automatically on its behalf. This deliberate aiming is the core idea of “active perception.”
  • Track / Player — a recording of a flight, plus the in-house viewer we use to replay it and watch what the drone did.

1. What I wanted

I wanted the drone to stop cheating physics.

In the old simulation the motion model was a convenience: I requested a velocity and the body adopted it on the same simulation tick, with the airframe staying dead level the whole time. It made early training tidy, but it taught the policy a world that does not exist. A policy trained on instant, level motion has no concept that you can’t stop on a dime, no concept that you must lean to move, and no concept that leaning takes time to build and time to bleed off. Drop that policy onto a real quadcopter and it would command an instant halt one meter from a wall and then drift straight into it, because the real airframe is still carrying its momentum.

You called this out directly, and you were right. So the goal for this entry was twofold:

  1. Give the drone a believable flight model — inertia, tilt-to-fly, and the natural overshoot-and-settle that comes with it.
  2. Hand the scanner over to the agent so that pointing the beam becomes a decision the policy is rewarded or punished for, not a free, automatic sweep.

2. What I tried

2.1 Real(istic) flight physics

I replaced the instant-velocity shortcut with a proper second-order motion model — the standard textbook description of how a quadcopter actually moves. To go forward, the drone now has to pitch forward first; the forward thrust is a consequence of the lean. Speed ramps up over time rather than snapping to its target, and the tilt angle itself behaves like a real control loop: it rushes toward the commanded angle, overshoots it, swings back through the other side, and gradually damps down to steady state.

The important honesty point: I did not invent the numbers. I took the established quadcopter model from the literature and reproduced its behavior, so the inertia and the wobble fall out of physics rather than out of a knob I twisted until it looked nice.

A concrete picture of the consequence: if the drone has built up real speed and a wall appears, it cannot cancel that speed in a single tick. It has to tilt back, fight its own momentum, and decelerate over several ticks — exactly the situation a real pilot or a real autopilot has to plan around. That constraint is now baked into every decision the agent makes.

2.2 A conscious scanner

Previously the long-range fan-scan swept on its own, like a lighthouse, regardless of what the drone needed to see. I changed it so the agent itself decides where and when to aim the TF-Luna beam. The short-range proximity sensors — the ones whose only job is “don’t clip the wall right next to you” — stay fully automatic and always on; you never want collision avoidance to depend on the policy remembering to look. But the long-range picture of the room is now something the drone must actively build.

The mechanism that makes this matter is a freshness rule on the drone’s internal map. If the agent hasn’t pointed the scanner toward a given direction for a while, that region of its “vision” goes stale and gets marked unknown. Flying into the unknown is discouraged. So there is constant pressure to keep sweeping the fan toward wherever the drone intends to go — not because a rule says “scan now,” but because failing to scan leaves blind spots that are costly to fly into. The desired behavior — look before you leap — emerges from that pressure rather than from a hand-coded scan schedule.

2.3 Make it visible in the Player

A wobble you can only read off a column of numbers is hard to trust. I added the drone’s tilt angle to the flight recording so the lean and the rock-back-and-forth are visible directly on the replayed track. Now I can watch the airframe pitch, overshoot, and settle, instead of squinting at telemetry.

3. What happened

Every test passed, and the recorded flight shows the new behavior clearly and exactly as the physics predicts.

Quantity Behavior on the recorded flight
Tilt (pitch) Ramps up to ≈ 9.6°, overshoots, swings back through negative, then settles — visible wobble
Forward speed Builds smoothly from ≈ 0.06 up to ≈ 0.53 instead of jumping instantly
Scanner Aimed by the agent’s own actions; stale directions correctly fall back to “unknown”

What I read off the track: the tilt doesn’t jump to its target and freeze — it climbs to about 9.6 degrees, sails past, dips to the other side, and then damps down. That little rock-back is the overshoot, and seeing it is how I know the second-order model is genuinely engaged. The speed curve tells the same story from the other direction: it eases in gradually, which is inertia made visible. Put together, this is finally a drone that flies rather than one that hops between grid cells.

4. Sources

  • Standard quadcopter flight-dynamics model from the published literature (reproduced, not invented) — the basis for the inertia and tilt-overshoot behavior.
  • Our own flight recordings and the Player viewer, used to confirm the tilt, wobble, and speed curves visually rather than only numerically.

5. What’s next

  • Swap the single test room for the real A.1 maps (7 scenes) so the policy meets varied, realistic indoor layouts instead of one toy space.
  • Add a colored semantic layer to the Player so the replay shows not just geometry but what each surface is.
  • Move training onto the GPU now that the dynamics and perception are honest enough to be worth the compute.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR