Hovering is the simplest RL task for a drone: stay at a target altitude (typically 1 m above ground) with zero horizontal velocity. The reward is straightforward — small per-step bonus for being within tolerance, penalty for deviation, large penalty for crashing.
This task isn’t where the interesting research lives — modern flight controllers do hovering perfectly without any learning. But it’s a useful pipeline smoke test: if your RL agent can’t learn to hover, something in the env / observation / action / reward stack is broken before more interesting tasks become tractable.
Current status: the hovering task exists as a debugging environment. We use it to verify that the policy training loop, the action publisher, and the simulated drone are correctly wired up. None of the 24 dev-log experiments target this task — they all run on obstacle avoidance (coverage in a cluttered map).
Typical reward
+0.1 per step within tolerance (e.g., ±10 cm altitude, ±0.1 m/s velocity)
-0.5 per step outside tolerance
-50 for crashing (ground or ceiling impact)
+10 terminal bonus for completing 30 seconds in tolerance
Where to go next
- Obstacle avoidance task — the actual research task
- Environments hub — sibling pages
- First PPO run — example of pipeline bring-up