orchestrator.py is the top-of-stack script that starts a full training (or evaluation) session: bring up Gazebo, start ArduPilot SITL, wait for MAVROS heartbeat, launch the RL policy node, watch for crashes, restart on failure, and tear everything down cleanly at the end. It’s the difference between “I can train for 12 hours” and “I babysit the terminal.”
Current status: placeholder. Today the orchestration lives in help_scripts/launch.sh (bash), which works fine for interactive use but doesn’t handle multi-day runs gracefully — no automatic restart on Gazebo crash, no checkpoint resumption logic. A Python orchestrator.py is on the backlog.
Planned responsibilities:
- Bring up Gazebo, ArduPilot SITL, MAVROS in the right order with proper inter-process wait conditions.
- Heartbeat watchdog: detect when MAVROS stops emitting and either retry or fail loudly.
- Checkpoint cadence: trigger model-save events at configurable intervals; resume from last checkpoint on restart.
- Clean teardown: SIGTERM in reverse order, capture final logs, archive
~/drone_media/sim/<session>/.
Why not just bash:
Bash is fine for the happy path. The unhappy paths — Gazebo segfaulting at minute 47, MAVROS losing connection mid-episode, a runaway training loop OOMing the host — are where you want structured exception handling, retry logic, and proper logging. That’s where a Python orchestrator earns its keep.
Where to go next
- Custom scripts hub — sibling utilities
- SITL bring-up — the thing being orchestrated
- Latest dev-log — when this lands, it’ll be there