Where we are at the start of H2
By end of June 2026, we have:
- Working PPO policy trained on
warehouse_v2, beating classical baselines on cluttered indoor coverage (see RL vs lawnmower). - Stable 3-rangefinder hardware stack — TF-Luna + VL53L0X + PMW3901 — at sub-$80 sensor cost.
- Full Gazebo Harmonic + ArduPilot SITL pipeline running end-to-end on a single host.
- One known out-of-distribution failure mode (
map_04open warehouse — see benchmark).
What we don’t have yet: a flying physical drone with the policy on it. The whole rest of this roadmap is about closing that gap.
Three workstreams
1. Sim-to-real (July–September)
The current policy works in sim. The bet for H2 is that with domain randomization and a noise-aware training pass, that policy transfers to the real drone with under 10% performance degradation. Two milestones:
- NEXT-A (noise-aware policy) — retrain with σ ∈ {0, 0.1, 0.2, 0.3} TF-Luna noise interleaved. Hypothesis: implicit regularization closes the
map_04gap. Target: ≥0.90 coverage on out-of-distribution maps. Status: in dev-log under TASK-RL-NEXT-A. - NEXT-MAP-AUG (map augmentation) — train across procedurally-generated indoor maps, not just
warehouse_v2. Target: stable performance across map families. This is the bigger fix and the slower workstream.
2. Robustness (August–October)
Real drones break. The current policy assumes all three rangefinders are healthy; real-world deployment will see TF-Luna occlusion, VL53L0X cross-talk, and PMW3901 dropout in low-texture environments. Robustness workstream:
- NEXT-NF (sensor-failure policy) — policy that explicitly handles “vl53l0x forward-failure” and similar partial-sensor-failure modes. Status: dev-log TASK-RL-NEXT-NF.
- NEXT-FRONTIER (frontier-weighted exploration) — combine learned policy with classical frontier exploration for better generalization to unknown spaces. Status: dev-log TASK-RL-NEXT-FRONTIER-MC-03.
3. Real-drone bring-up (September–December)
The physical platform. Three milestones:
- Drone assembly — sourcing, soldering, calibration. Frame is a 5-inch indoor quad; ESC/motor selection driven by the BOM. Target: flyable manual-control drone by end of September.
- MAVROS + companion-pc integration — Jetson Nano running the policy on-device, communicating with ArduPilot over MAVROS. See MAVROS theory. Target: first closed-loop flight (policy commanding setpoints, drone executing) by end of October.
- Indoor flight campaigns — real-drone benchmark runs in a physical mockup of
warehouse_v2. Compare against simulation results and the classical baseline. Target: first benchmark paper or technical report by end of December.
What’s explicitly not in H2
- Outdoor flight. Different problem, different platform, different regulations. Not in scope.
- Multi-drone coordination. Interesting future direction but premature until single-drone behavior is solid.
- Visual SLAM. Considered and rejected for now — adds compute weight without solving the indoor obstacle-avoidance problem better than the rangefinder stack already does.
- Investor-facing demos. We’ll have something demo-able by end of Q4, but the focus is on the engineering being real, not on running early demos that look better than they work.
Risks and what could go wrong
- Sim-to-real gap is larger than 10%. Recovery: more aggressive domain randomization, possibly real-data fine-tuning. Cost: 2-3 months delay.
- Hardware bring-up takes longer than September. Indoor drone build always has unknowns (motor balance, frame resonance, ESC programming). Recovery: parallelize with NEXT-A; the sim work can absorb the slack.
- Out-of-distribution generalization doesn’t close. Recovery: narrow the deployment scope to maps that look like training data. Less ambitious but still useful.
Where to go next
- Latest dev-log — current week-by-week progress against this roadmap
- RL framework — the technical substrate the H2 work builds on
- Hardware overview — what the physical drone will be