claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

29 — Alignment-v1 Acceptance and Takeoff Fixes

How we fixed post-reboot takeoff crashes, accepted the N=6 aligned policy, and wired the distance-sensor chain through to the flight controller.

stablesimulationupdated 2026-06-16T00:00:00.000ZClaudeDroneSimulationDevLog

What this is: A simulation dev-log entry covering a single working session: chasing down why the drone kept crashing on takeoff after a machine reboot, accepting the first aligned navigation policy through a formal quality gate, and proving that our distance-sensor data actually reaches the flight controller.

Why it’s here: It’s a candid look at how a lot of robotics work is unglamorous plumbing — settling timers, hover targets, and parameter files — and how a careful acceptance test catches problems before they reach a real drone.

Date: 2026-06-08 Ticket: Alignment-v1 acceptance + takeoff fixes (branch v2_takeoff_plus_model)


Glossary

A few terms appear throughout, so let me define them in plain English first.

  • Acceptance test — a pass/fail check we run before we trust a piece of software. Think of it like a driving test: it doesn’t matter how good the car feels, it either clears the course within the rules or it doesn’t. Ours measures a specific failure (the drone telling itself to move but not actually going anywhere) and demands it stay under a strict threshold.
  • Alignment v1 — “alignment” here means making the simulator’s behavior match the exact conditions the navigation model was trained under. If the model learned to fly in one world and we test it in a subtly different one, it misbehaves — like a pianist trained on one keyboard suddenly handed one where the keys are slightly the wrong width. “v1” is our first complete pass at this matching.
  • Takeoff (in our stack) — the drone arms its motors, lifts off, and climbs to a hold altitude. It sounds trivial, but it’s where a surprising number of bugs surface, because everything (sensors, motor controllers, timing) has to be ready at exactly the same moment.
  • SITL — Software In The Loop: the real flight-controller firmware running as a program on a PC instead of on physical hardware. It thinks it’s flying a real drone.
  • EKF — the Extended Kalman Filter, the math the flight controller uses to fuse sensor readings into a single confident estimate of “where am I and which way am I pointing.” It needs a moment after startup to converge.
  • no-travel — our nickname for a specific failure: the policy issues a “move forward” command, but the drone doesn’t actually translate to a new position. Lots of these means the drone is spinning its wheels (figuratively).

1. What we wanted

This session ran right after a reboot of the GUI machine, and it had a clear chain of goals:

  1. Finish a navigation test around a central pillar that was already in progress.
  2. Fix takeoff, which had suddenly started crashing every single run after the reboot.
  3. Receive and validate the first aligned export of our six-sensor navigation policy.
  4. Run that policy through Aleks’s formal acceptance gate and decide whether it’s good enough to ship.

A nice property of this whole session: everything was fixable on the simulation side — bridge code, timers, and configuration. No model retraining was needed, which kept the iteration loop fast.

2. What we did

2.1 Two takeoff fixes (verified live)

After the reboot, three out of three runs crashed during takeoff with an attitude-error abort at roughly 0.2–0.9 m off the ground. It happened in both the GUI and the headless setups, and with or without wiping the firmware’s saved settings — which ruled out the usual graphics and persistence suspects. So I went looking deeper.

Two root causes, two fixes:

  1. An EKF settle gate. Gazebo’s simulated IMU comes up more slowly after a reboot. We had configured the firmware to skip its arming checks and disarm-delay for speed, which meant the drone could arm before the EKF had converged — its gyros still inconsistent — and then immediately tip into an attitude-error crash. The fix is simple and humble: wait a fixed settling period (about 8 seconds) after entering guided mode before arming. Give the math time to agree with itself.

  2. A hold-in-place hover target. This one was sneaky. The takeoff controller’s hover target was hardcoded to coordinates (0, 0) — the world origin. But the pillar obstacle is spawned offset from center, which meant the drone tried to fly toward the origin during takeoff, banked over to nearly 47 degrees, and sank from 2.1 m back down to 0.6 m — sometimes straight into the obstacle. The fix: read the actual takeoff x,y from odometry and hold that spot. This cured both “won’t hold altitude” and “drifts into the obstacle” in one stroke.

  3. A hover stability monitor. As a safety backstop, a monitor watches a tight altitude band and a descent-rate limit, and refuses to hand a falling drone back to the navigation bridge. During the crash investigation both safeguards fired live, which was reassuring — the net caught the drone.

All three were validated across six test takeoffs.

2.2 Pillar navigation test — passed

With takeoff stabilized, the in-progress pillar test completed cleanly: mission complete, coverage around 0.95, a clean on-axis track, and the central column fully circumnavigated. We saved image and animation artifacts of the flight path.

Pillar navigation track

2.3 A transparent ceiling and a doorway audit

Aleks asked for a transparent ceiling (a low-opacity lid at 4 m) added to all the reinforcement-learning rooms. It’s visually helpful and behaviorally neutral — it doesn’t change how the policy navigates.

While in the world files, I audited doorway widths and found a mismatch: one two-chamber world had an actual gap of 1.0 m, not the 1.2 m the documentation claimed. I exposed doorway_width as a real, validated world-configuration field so the truth lives in the data, not in a comment. (Backlog note: doorways narrower than ~1.4 m are effectively impassable once the six-sensor policy keeps a 1.0 m standoff, so wider doorways are coming.)

2.4 The N=6 aligned export — replay caught two gaps

We received the first aligned export of the six-sensor policy. Before deploying it, a parity replay test — which compares what the simulator feeds the model against what the model expects, observation by observation — caught two mismatches:

  1. The frontier observation was using raw sensor data instead of the filtered version the protocol specifies. The counts matched bit-for-bit (44 of 44) only with raw data, so the retrain hadn’t actually enabled the filter. The good news: raw is in distribution for the bridge, so it’s safe.

  2. The action mask was margin-aware in training, but the bridge was returning “all clear.” Here 37 of 44 cases diverged. The model was trained to respect a safety margin near walls; the bridge wasn’t mirroring that. This is exactly the kind of silent gap that a real flight would expose at the worst possible moment, so catching it in replay was a win. (This echoes an earlier lesson with our down-facing rangefinder — the simulator must mirror sensor handling precisely, not approximately.)

2.5 Acceptance test against the gate

The gate Aleks set: on an empty 6×6 m world, no-travel events must stay at or below 2%.

  • Run 1 (no extra safeguard): mission complete, but no-travel came in at 6.5% — a fail. Notably, all the failures clustered when the front sensor read 0.67–0.68 m, just under the executor’s 0.70 m margin. Still, compared to the previous-generation policy’s 25%, this was already a ~74% reduction.

  • A sensor-driven safeguard was added: before the model predicts, mask the “advance into frontier” action whenever the front sensor reads below the executor’s effective margin — and crucially, that margin is read from the active mode, not hardcoded.

  • Run 2 (with the safeguard): mission complete, no-travel down to 5.26% — still technically a fail, but the safeguard demonstrably worked, firing at least eleven times and converting would-be stalls near walls into rotations. The handful that leaked through were a timing gap: the model checks the sensor and sees clearance (≥0.70 m), then by the time the move executes the drone has crept to 0.65 m.

The root of the remainder: the executor’s advance margin (0.70 m, from the explore mode) sits above the standoff distance the policy was trained with (0.60 m). In the band between those two values the model considers advancing valid (it learned at 0.60), but the executor refuses — producing a no-travel. A safeguard set at 0.70 can’t fix a conflict that is 0.70.

Aleks’s decision: accept the 5.26%. It’s five times better than the previous generation, and the complete fix belongs in the next retrain (three coordinated changes folded into one training pass). The Alignment-v1 sprint is closed.

2.6 Corners — the honest weak spot

In the un-safeguarded run, 63% of the drone’s actions were rotations, plus two near-livelocks at walls (a heading-change feasibility backstop rescued both). The mechanism is instructive: during training the environment itself masked the advance action near walls, so the model never had to learn to actively turn away from them — the mask did that job. The bridge doesn’t mirror that mask, so the model dithers at walls and corners. The safeguard treats the symptom (rotate instead of stall) but not the cause. Aleks floated an idea (not yet implemented) to reward clean corner turns directly, which would address the cause rather than the symptom.

2.7 Distance-sensor end-to-end: proving the chain

A separate thread: an end-to-end check of the distance-sensor pipeline, from the simulated sensors all the way to the flight controller.

The transport side worked:

  • The perimeter sensor topic was live, reporting six channels at their max range (the drone sat centered, walls beyond range).
  • The bridge republished readings as proper Range messages at very close to 10 Hz, all valid.
  • The flight controller’s distance-sensor plugin was active across all six channels and forwarding the data onward as standard MAVLink distance messages.

The consumption side had a gap: the flight controller’s proximity feature was switched off, and its rangefinder type was set to the SITL analog default rather than the MAVLink type. So the data arrived, but the firmware ignored it — the parameter file simply didn’t contain the proximity and rangefinder configuration.

So I added the missing parameters to the indoor config: enable MAVLink proximity (forward orientation) and set the MAVLink rangefinder type. There’s an important caveat I flagged: the first rangefinder slot is already used by the default downward rangefinder that the simulated airframe uses for altitude, so changing its type risks breaking height estimation across every benchmark run. The down-facing sensor likely belongs in a second slot. I wrote out a verification plan (wipe saved settings first, confirm the parameter names weren’t silently ignored, confirm the controller receives the data, and — most important — confirm takeoff altitude still holds).

Finally, a dedicated forwarder node was wired in and verified: the six perimeter channels plus the down-facing laser flow through to the controller as MAVLink distance messages. On the empty 6×6 world it armed without a “No Data” pre-arm complaint, climbed to 2.0 m cleanly, and the channels reported sane ranges. End-to-end, confirmed.

3. Results

  • Takeoff is fixed and stable: 6/6 test takeoffs with zero crashes, both GUI and headless.
  • The pillar test passed with ~0.95 coverage.
  • The N=6 aligned policy was accepted at 5.26% no-travel — a 5× improvement over the prior generation — and the Alignment-v1 sprint is closed.
  • A later overnight test of the masking fix on the next-generation policy hit 0% no-travel on the acceptance world (mission complete, ~0.91 coverage, a healthy non-degenerate flight), confirming the fix direction. That export stays in testing only, not production, per Aleks. The current production policy remains the accepted N=6 v1.
  • The full distance-sensor chain — simulated sensor to flight controller — is wired and verified end-to-end.

The no-travel progression across the work tells the story cleanly: previous generation 25% → N=6 v1 6.5% → with safeguard 5.3% → masking fix 0%.

N=6 acceptance track on empty 6×6

4. Sources

  • Branch v2_takeoff_plus_model, simulation repository.
  • Flight-path artifacts: track_pillar_v15c, track_n6_empty6x6, track_n6b_empty6x6, track_v2_empty6x6 (image + animation), in the simulation media store.
  • The acceptance pipeline that computes no-travel, including for short runs.

5. What’s next

  • Fold the three coordinated alignment fixes into the next training pass (the wall-standoff masking, the frontier filter, and a sensor-aware action mask).
  • Add wider doorways (≥1.4 m) to the worlds so the six-sensor standoff can fit through.
  • Validate the new proximity and rangefinder parameters in a benchmark batch, with altitude sanity as the first checkpoint — and roll back the rangefinder change immediately if it disturbs height estimation.
  • Explore rewarding clean corner turns to fix the wall-dithering at its cause rather than masking the symptom.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR