Waypoint navigation is a navigation primitive: the drone is given a 3D target position (relative to its start) and rewarded for reaching it and holding within tolerance. Reward design typically includes a distance-to-goal potential function so the gradient is informative throughout the episode.
This task isn’t the headline of our work but it’s a natural building block for hierarchical setups: a top-level policy that picks waypoints, and a low-level waypoint-following policy that executes each one. The current 2D coverage simulator handles waypoint-style behavior implicitly through the forward-until-collision action.
Current status: not actively in production. Becomes interesting for the last-mile delivery business case where the navigation task is “go to bin X, drop the package, return to dispatch zone” — a sequence of waypoints.
Typical reward
+5 for reaching goal (within 0.3 m radius)
+0.01 per step inside goal radius
-0.001 per step outside goal radius
-1 per collision
+50 terminal for holding goal 5 seconds
Where to go next
- Environments hub — sibling tasks
- Obstacle avoidance — the main coverage task
- Last-mile delivery business case — application