claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

Task: obstacle avoidance

Navigate cluttered indoor environments without collisions while maximizing map coverage. The main RL task.

stabledoc-seoupdated 2026-05-11T00:00:00.000ZClaudeDrone

Obstacle avoidance is the central reinforcement-learning task in the claudeDrone framework. The drone is spawned in a cluttered, maze-like indoor environment, navigates using onboard rangefinder data and a memory of where it has already been, and must explore as much of the space as possible without colliding with walls or obstacles. The objective is to maximize the fraction of the map covered within a bounded episode.

The simulation runs on Gazebo with ROS2 Jazzy, using cluttered warehouse- and maze-style worlds. See Gazebo worlds.

What the agent observes

The policy receives a structured observation combining three sources:

  • Rangefinder distances — readings from a ring of perimeter distance sensors plus a steerable sweeping sensor, giving the agent a sense of nearby obstacles in several directions. Distances are normalized to a common scale.
  • Sensor orientation — the current pointing angle of the steerable sweep sensor.
  • Visited map — a coarse occupancy-style grid marking which cells the drone has already visited. This acts as the policy’s spatial memory, helping it avoid revisiting explored areas.

What the agent can do

The agent chooses from a small, discrete set of high-level moves: stepping into an adjacent cell, rotating in place by a fixed increment, performing a scan with the steerable sensor, and a compound “move forward until an obstacle is reached” action. The compound forward action was a notable step in the project’s development, since it lets the agent cover open corridors efficiently in a single decision rather than one cell at a time.

How behavior is rewarded

The reward is shaped to encourage thorough, safe exploration. In conceptual terms, the agent earns reward for discovering previously unvisited cells, receives a smaller incentive for actively scanning its surroundings, and is penalized for collisions. A small per-step cost discourages aimless wandering and pushes the policy toward efficient routes. Reaching near-complete coverage of the map ends the episode with a substantial bonus. The balance between these incentives steers the agent toward covering ground quickly while keeping collisions rare.

How it performs

The headline comparison for this task pits the learned policy against a classical sweeping (“lawnmower”) baseline. The learned policy reaches useful coverage quickly and clearly outperforms the baseline over short horizons, where systematic sweeping has not yet had time to pay off. Over very long horizons the exhaustive baseline eventually edges ahead on total coverage, while the learned policy remains far more sample-efficient early on. See RL vs lawnmower benchmark for the full discussion.

Where to go next

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR