Obstacle avoidance is the central reinforcement-learning task in the claudeDrone framework. The drone is spawned in a cluttered, maze-like indoor environment, navigates using onboard rangefinder data and a memory of where it has already been, and must explore as much of the space as possible without colliding with walls or obstacles. The objective is to maximize the fraction of the map covered within a bounded episode.
The simulation runs on Gazebo with ROS2 Jazzy, using cluttered warehouse- and maze-style worlds. See Gazebo worlds.
What the agent observes
The policy receives a structured observation combining three sources:
- Rangefinder distances — readings from a ring of perimeter distance sensors plus a steerable sweeping sensor, giving the agent a sense of nearby obstacles in several directions. Distances are normalized to a common scale.
- Sensor orientation — the current pointing angle of the steerable sweep sensor.
- Visited map — a coarse occupancy-style grid marking which cells the drone has already visited. This acts as the policy’s spatial memory, helping it avoid revisiting explored areas.
What the agent can do
The agent chooses from a small, discrete set of high-level moves: stepping into an adjacent cell, rotating in place by a fixed increment, performing a scan with the steerable sensor, and a compound “move forward until an obstacle is reached” action. The compound forward action was a notable step in the project’s development, since it lets the agent cover open corridors efficiently in a single decision rather than one cell at a time.
How behavior is rewarded
The reward is shaped to encourage thorough, safe exploration. In conceptual terms, the agent earns reward for discovering previously unvisited cells, receives a smaller incentive for actively scanning its surroundings, and is penalized for collisions. A small per-step cost discourages aimless wandering and pushes the policy toward efficient routes. Reaching near-complete coverage of the map ends the episode with a substantial bonus. The balance between these incentives steers the agent toward covering ground quickly while keeping collisions rare.
How it performs
The headline comparison for this task pits the learned policy against a classical sweeping (“lawnmower”) baseline. The learned policy reaches useful coverage quickly and clearly outperforms the baseline over short horizons, where systematic sweeping has not yet had time to pay off. Over very long horizons the exhaustive baseline eventually edges ahead on total coverage, while the learned policy remains far more sample-efficient early on. See RL vs lawnmower benchmark for the full discussion.
Where to go next
- RL vs lawnmower benchmark — coverage comparison against the classical baseline
- Environments hub — sibling tasks and related environments
- Gazebo worlds — the worlds this task runs in