The deployment step is where a trained PPO checkpoint stops being a .pt file on the dev host and becomes a process running on the drone’s companion PC, consuming sensor topics and producing action commands at 10 Hz. This page covers the practical mechanics of that handoff: how to package the checkpoint, what the inference loop looks like, and how to validate that the policy on the drone is the same policy that you evaluated in simulation.
Current status: placeholder. Real-drone inference is on the H2 2026 roadmap — specifically the September-October “MAVROS + companion-pc integration” milestone. Today, the inference pipeline runs only in simulation, on the same host that did the training. The architecture below is what we plan to ship.
Planned deployment flow
- Checkpoint export. Trained PPO model exported to TorchScript (
.pt) on the dev host. Includes the policy network only — no value function, no optimizer state, no replay buffers. Target size: < 50 MB. - rsync to companion PC. Copy the checkpoint to the Jetson under
/opt/claudedrone/checkpoints/<exp-id>.pt. The Jetson has no internet access in flight — pre-stage everything. - Inference node start. ROS 2 node loads the checkpoint, subscribes to the observation topic (
state_estimatoroutput), publishes actions toaction_publisher. - Smoke test on the ground. Drone disarmed; verify the inference loop produces sensible actions when you wave a hand in front of the rangefinders. Catches the “wrong observation schema on the deployed checkpoint” class of bug before motors start.
- Hover sanity check. Arm, GUIDED, take off manually to 1 m. Switch to RL-control mode. Hover for 30 s. If the drone doesn’t immediately try to fly into something, the deployment is probably correct.
Inference loop sketch
class RLInferenceNode(Node):
def __init__(self):
super().__init__('rl_inference')
self.policy = torch.jit.load('/opt/claudedrone/checkpoints/exp_7.pt')
self.policy.eval()
self.sub = self.create_subscription(
DroneState, '/drone/state/observation', self._on_state, 10)
self.pub = self.create_publisher(
Float32MultiArray, '/policy/action', 10)
def _on_state(self, msg):
obs = self._build_observation(msg)
with torch.no_grad():
action = self.policy(obs).numpy()
self.pub.publish(Float32MultiArray(data=action.tolist()))
A real implementation will add observation-buffer staleness detection (drop the inference call if the observation is more than ~200 ms old), action smoothing across consecutive timesteps, and a per-axis clipping layer that prevents wildly out-of-distribution actions.
What to verify before any real-drone deployment
- Observation schema parity. The simulated training observation must match the deployed observation, exactly. A reordered field will pass type checks and break flight.
- Sensor health gating. The policy was trained on healthy sensors. If a sensor goes invalid mid-flight, the inference node should fall back to a safe action (hover, or hand off to a classical fallback controller) rather than feed garbage into the policy.
- Latency budget. The simulation observation arrives at the policy ~5 ms after the underlying sensor reading. The real-drone observation may arrive 50-100 ms later (companion-PC IO + MAVROS bridging). The policy was trained assuming the simulated latency — if the real latency is significantly worse, performance will degrade.
Where to go next
- RL framework — what the policy is and how it was trained
action_publishernode — the consumer of this node’s output- Operations hub — sibling operational pages
- Roadmap H2 2026 — when this actually ships