claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

Teaching the drone to fly with RL

Reinforcement learning stack: tasks, observations, actions, rewards, algorithms, metrics.

draftdoc-seoupdated 2026-05-11T00:00:00.000ZClaudeDrone

The RL stack is the reason claudeDrone exists. Classical autopilots can already hover, follow waypoints, and avoid obstacles — but they require careful per-environment tuning and degrade fast when the world doesn’t match the assumptions a human encoded in the gain table. The bet here is that a policy trained in simulation can replace those hand-coded behaviors with something that generalizes better, especially in cluttered indoor spaces where GPS is gone and the obstacle distribution looks nothing like the open-air datasets most autopilots assume.

Concretely, the stack is PPO (with SAC as a fallback for exploration-heavy tasks), trained in warehouse_v2 on observation vectors built from TF-Luna lidars and a BNO085 IMU. The action space is normalized velocity — we don’t have the policy publish raw PWM because the sim-to-real gap on PWM dynamics is brutal. Rewards are a layered combination of distance-to-goal, energy efficiency, and crash penalties, tuned so the policy prefers smooth trajectories without explicitly being told to.

This section has six sub-areas. Environments defines the RL tasks (hover, waypoint, obstacle avoidance). Observation space and Action space define what the policy sees and does. Reward functions is where most of the engineering judgment lives — reward design is the single biggest lever on what behavior the policy learns. Algorithms covers the PPO and SAC implementations and the hyperparameters that worked best for us. Training metrics explains what we look at in TensorBoard and how to read the standard learning curves.

If you’re stepping into RL on this project for the first time, start with reward functions — not algorithms. The choice of PPO vs SAC matters less than whether the reward gradient actually points toward useful behavior, and most of our early experiments failed because the reward design was off, not because the optimizer was. Once you’ve internalized the reward design, PPO hyperparameters covers the knobs that actually move performance.

Contents

Auto-generated from child _index.md files and entries during build (update-indexes.mjs).

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR