What this is: A working note about the day I sat down and wrote out, in black and white, exactly what data flows in and out of one specific piece of our drone’s brain — the navigator node — before I let it learn anything. Why it’s here: In robotics, the most expensive mistakes are the silent ones: two modules quietly disagreeing about what “left” means. This entry is the story of nailing down those agreements first, so training doesn’t bake in a bug.
Date: 2026-06-16 Ticket: Sprint A.1 — navigator RL node (T04)
Glossary
- Navigator node — one software component on the drone whose only job is to decide where to fly next. Think of it as the part of your brain that picks a direction in a dark room; it does not run your legs or your eyes, it just chooses. Here it is the learning component, and everything else (sensors, maps, mission goals) feeds into it.
- Interface contract — a written, frozen agreement about the exact shape and meaning of data passed between two modules. Like a power socket standard: as long as both the plug and the wall agree on the shape, you can swap either side without anything melting. Skip this, and one module sends centimeters while the other reads meters.
- A.1 sprint — the very first stage of a longer plan. “A.1” is just a label, like “Chapter 1.” The goal of this stage is the most basic version of flight that still teaches us something, using a simplified world rather than the full physics simulator.
- PPO — Proximal Policy Optimization, the reinforcement-learning algorithm doing the actual learning. You can picture it as a cautious trainee that improves its strategy in small, careful steps rather than wild leaps.
- Observation space — the complete list of everything the navigator is allowed to “see” on each tick: a small local map plus a vector of numbers (sensor readings, position, target offset).
- Action space — the complete menu of moves the navigator is allowed to choose from. Ours is discrete: a fixed list of named choices (forward, rotate, hover, …) rather than a continuous steering dial.
- Parity point — a specific detail where my simplified training world and the full simulator must agree exactly. I number them (P1, P2, …) so nothing slips through. “Parity” here means “identical on both sides.”
- ROS2 (Jazzy) — the robotics middleware our nodes talk over. REP-103 below is its convention for which way the coordinate axes point.
1. What I wanted to do
The bigger A.1 plan needs six cooperating software nodes. My slice is exactly one of them: the navigator RL node. The other five — map extraction, semantic zoning, waypoint publishing, episode management, and scene setup — belong to the Simulation agent.
Here is the catch that drove this whole entry: my node consumes the outputs of those five. If their data shape or meaning is even slightly off from what I assume, my training run learns garbage and I won’t find out until much later. So before writing a single line of training code, I wanted the contract — observations in, actions out — written down and agreed.
A few framing decisions were settled in conversation with Aleks:
- Actions are discrete. Each chosen action is held for a short window and translated into a steady velocity command, not a teleport between grid cells. So it flies smoothly even though the decisions are a small menu.
- The training substrate is a 2D abstraction. The full Gazebo simulator is reserved for validation; learning happens in a lighter world that runs far faster.
- The privileged “Teacher” setup. For A.1 the navigator is allowed to see ground-truth information it won’t have in the real world. We strip those privileges in a later stage.
I also carried in a scar from an earlier sensor-alignment bug (a 90-degree mismatch on a range sensor), which is precisely why I refused to start training before pinning the parity points.
2. What I tried — the contract itself
2.1 Observation space — a dictionary of two parts
I settled the observation as a dictionary with two members, which lets the policy use a small image network for the map and a simple network for the rest:
map— a 2-channel grid, a 5m × 5m cut-out centered on the drone.vec— a 52-number vector of everything else, in a strict, fixed order.
The two map channels:
| ch | field | meaning |
|---|---|---|
| 0 | ground-truth map patch | free vs. unknown vs. wall, as a continuous shade |
| 1 | local visit count | how often each nearby cell has been visited |
One subtlety I flagged immediately (parity point P1): the simulator’s master map and my local patch use different resolutions. The fix is to resample the simulator’s map into my patch’s resolution, and on my 2D substrate I simply build the map at the patch resolution directly. When we later bridge to the simulator, that bridge must mirror it.
The 52-number vector, in order:
| idx | field | size | note |
|---|---|---|---|
| 0…15 | semantic zone | 16 | which kind of space the drone is in, plus a confidence value |
| 16…18 | waypoint offset (dx,dy,dz) | 3 | direction to the current target, clamped |
| 19…21 | position (x,y,z) | 3 | meters from a scene origin |
| 22 | yaw (heading) | 1 | scaled to [-1, 1] |
| 23…25 | velocity (vx,vy,vz) | 3 | clamped |
| 26…31 | short-range sensors ×6 | 6 | normalized by max range |
| 32 | downward range | 1 | height, normalized |
| 33…51 | forward range scan | 19 | a fan of beams across a half-circle, normalized |
That sums to exactly 52, and writing it out this explicitly is the whole point — every consumer now knows that index 22 is heading and not, say, the third velocity.
The other parity points I recorded before training:
- P2 — the exact ordering and mounting angles of the six short-range sensors.
- P3 — the forward scan really is 19 beams across the half-circle, which differs from an older 12-sector layout I’d used. Confirm beam count and spacing with sim.
- P4 — how the 16-slot semantic vector is laid out (which slot carries confidence vs. an id).
- P5 — how the “scene origin” is fixed when the drone spawns at a random spot.
2.2 Action space — a menu of 11 moves
Actions are a discrete list of eleven. Discrete on purpose: it’s easier to debug, and each move maps cleanly onto a mission primitive later.
| ID | action | command (body frame) |
|---|---|---|
| 0 | FORWARD | forward along heading |
| 1 | BACKWARD | backward |
| 2 | STRAFE_LEFT | crab left, no turn |
| 3 | STRAFE_RIGHT | crab right, no turn |
| 4 | ROTATE_CW | turn clockwise |
| 5 | ROTATE_CCW | turn counter-clockwise |
| 6 | UP | climb |
| 7 | DOWN | descend |
| 8 | HOVER | hold position |
| 9 | SCAN_FAN | sweep the scanner across a fan |
| 10 | SCAN_PRECISE | stop, stabilize, take a precise reading |
Each chosen action is held for a few consecutive ticks, so the drone re-decides a few times per second while velocity streams continuously underneath — discrete decisions, smooth flight.
The single most dangerous detail lives here, parity point P6 (the sign of sideways velocity). The original spec document implied one sign convention for left/right strafing; the actual node uses the ROS2 (Jazzy) body-frame convention (REP-103, y-axis points left). I had initially copied the wrong signs from the document. That is exactly the class of silent disagreement this whole exercise exists to catch — getting it backwards would have the drone strafe the opposite way on command. The table above is corrected, and the authoritative signed parity table now lives in the interface dev-log.
A second action-side note (P7): on the 2D substrate there’s no physical servo to sweep, so I model the scan actions as a ray-cast refresh of the scan buffer. The real scan semantics for Gazebo validation are pinned separately.
2.3 Reward — kept multiplicative, watched closely
The reward is a multiplicative shape: a base term, scaled by per-zone multipliers and a visit-based factor, minus penalties (plus combo bonuses). I’m deliberately not reproducing the weights or formula here; the full table lives in the training config.
What matters for this log is the risk thinking I applied up front, because a multiplicative reward with many zones and combo bonuses is a textbook invitation to reward-hacking — the agent optimizing the score in ways that miss the actual goal. So before training, not after a failed run, I flagged things to trace:
- Whether combo chains (e.g. a room-transition bonus) could be farmed by looping between two zone types — I need to check those loops are even reachable.
- Whether the new-cells term and the visit factor double-count the same behavior, since one already zeroes out on revisits while the other penalizes them.
- Whether the scan-dependent combos are reachable by a real policy rather than sitting inert.
- The balance between the per-step time cost and the exploration reward, so the drone neither stalls nor sprints recklessly.
- Sharp zone thresholds creating cliffs in the reward, which I’ll want to sample densely around when tuning.
3. What happened
The contract is frozen. Observations, actions, and reward shape are written down with every index and sign accounted for, and the seven parity points (P1–P7) are recorded as explicit agreements the simulator side must mirror bit-for-bit when we bridge later.
The acceptance targets are also set. We track these on TensorBoard: episode reward, collision rate, waypoint success rate, coverage percentage, zone stagnation, episode length, and policy entropy. The gates we hold the trained policy to:
| metric | gate |
|---|---|
| collision rate | < 5% |
| waypoint success rate | > 80% |
| coverage | > 60% |
| policy entropy | not below 0.05 (to avoid policy collapse) |
Missions are defined as a set of YAML files across four difficulty levels — from simple waypoint following up to topology and reconnaissance — and the policy only advances to the next level once it clears an 80% success rate. A1’s env phase starts at level one: a straight-corridor mission.
The environment code itself is being written in a separate phase, on Aleks’s go. This entry’s deliverable was the contract, and that’s done.
4. Sources
- The two A.1 specification documents (observation/action/reward definitions and the master plan).
- A donor codebase from an earlier flight-RL effort that I’m reusing piece by piece:
| component | reused for |
|---|---|
| sensor ray-casting | the six short-range sensors, the 19-beam scan, and scan actions |
| frontier detector | the frontier zone and a waypoint fallback |
| privileged-data plumbing | the A.1 privileged Teacher setup |
| domain-randomization hooks | a later stage where privileges are removed |
| occupancy/observed grid | the 2D-substrate ground-truth map patch |
| training harness | extended with an A.1 environment branch |
- The interface dev-log, which now holds the frozen, authoritative signed parity table.
5. What’s next
- On Aleks’s go, finish the A.1 2D-abstraction environment.
- Wire the A.1 environment branch into the training harness.
- Run a smoke test: reset, step, observation shapes, and reward signs, then the first straight-corridor mission.
- Hand the recorded P1–P7 parity values to the Simulation side before Gazebo validation.