claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

48 — Reward Revision: A Minimal Reward Finally Solves Scene S1

Stripping the reward down to its bare essentials cured a self-crashing drone agent and let PPO solve scene S1 at 100% waypoint success, 0% collisions.

stablerl-labupdated 2026-06-16T00:00:00.000ZClaudeDroneRLDevLog

What this is: An engineering-journal entry about a small but decisive change to how we reward the drone-flying agent during training, and how that change finally got it to fly cleanly through our first test scene.

Why it’s here: It’s a clean illustration of a classic reinforcement-learning trap — an agent that learns to crash on purpose — and how going back to the simplest possible incentive fixed it.

Date: 2026-06-16 Ticket: A.1 (FlightRL-v3)


Glossary

A few plain-English definitions before we dive in, with everyday analogies so the rest reads smoothly even if you’ve never touched reinforcement learning.

  • Reward — a single number we hand the agent after each tiny step of flight, telling it “that was good” (positive) or “that was bad” (negative). The agent’s whole job is to rack up as much total reward as it can over an episode. Think of it like points in a video game: the agent doesn’t know what we want, it only chases points, so the points have to be designed to mean “fly well.”
  • Scene S1 — the easiest level in our training curriculum: a single straight corridor with a few waypoints the drone has to fly through in order. It’s the “tutorial level” before we move on to harder rooms and recon-style scenes later.
  • Solved — for us, “solved” means the trained agent reliably reaches every waypoint and never hits a wall during evaluation. Concretely: 100% waypoint success, 0% collisions.
  • PPO (Proximal Policy Optimization) — the learning algorithm we use to train the flying policy. The mental image: PPO nudges the agent’s behavior toward whatever has been earning more reward, but only in small, cautious steps so it never lurches into something wildly different and self-destructs. It’s the workhorse of modern robotics RL.
  • Waypoint — a target point in space the drone must pass through. A corridor mission is just a chain of these, like checkpoints in a race.
  • Episode — one full attempt: the drone spawns, flies until it either finishes the mission, crashes, or runs out of time, then we reset. Training is millions of these attempts in fast-forward simulation.

1. What I wanted

Coming out of dev-log 47, I had a stubborn failure on my hands. The agent on scene S1 wasn’t slowly getting better and then stalling — it was doing something far more annoying: it was crashing almost immediately, on purpose, and seemed perfectly happy about it. My goal for this session was narrow and concrete: figure out why, and get S1 to “solved” (100% waypoints, 0% collisions) at the 100k-step training budget I use for cheap smoke tests.

The reason this matters: S1 is the foundation. If the agent can’t learn to fly a simple straight corridor, there’s no point reaching for the harder scenes. And a classical (non-learning) planner already flies S1 in about 198 steps — that’s my benchmark. I wanted the RL agent to at least match it.

2. What I tried

I treated this as a quick hypothesis-elimination loop. A full 100k-step smoke run finishes in about 27 seconds, so I could test ideas almost as fast as I could think of them.

  • Hypothesis: maybe the incentive is just a bit too harsh. I softened several of the per-step nudges, leaned a little on staying alive, and re-ran. No change — the agent still bolted for a wall.
  • Hypothesis: maybe the corridor is too narrow and the agent feels trapped. To test this cleanly, I swapped in a wide-open hall instead of the tight corridor — same incentives, just vastly more room to maneuver. If narrowness were the problem, the agent should suddenly behave. It didn’t. It crashed just as fast in the open hall. That was the tell: the problem was not the geometry of the scene. It was the reward itself.
  • Hypothesis: strip the reward to the bare minimum. This was the one that worked. I cut the incentive down to only the essentials — reach the waypoints, finish the mission, and don’t crash — and switched off every secondary “polish” signal that had been nudging the agent on every single step.

3. What happened

The minimal reward solved it outright. On the S1 corridor evaluation the agent hit 100% waypoint success with 0% collisions, reaching all of the mission’s waypoints and flying the route in about 178 steps — actually a touch tighter than the ~198-step classical baseline. In one 100k-step training budget, the learning agent caught up to the hand-built planner on S1.

The training curves told the same story plainly. The average reward per episode went from clearly negative to solidly positive, and episodes got much longer — from a handful of steps (crash and reset) to a couple hundred steps of actual flying. That length jump alone is the signature of the cure: the agent had stopped ending its own episodes early.

Why was it crashing on purpose? This is the part worth remembering. I had been handing the agent lots of small per-step discouragements — gentle “be smoother,” “use less effort,” “behave” nudges applied on every single timestep. Early in training the agent’s actions are very noisy and exploratory, so on most steps it wasn’t yet earning enough good signal to outweigh that steady drip of small discouragements. The net effect: every step of life felt slightly bad. And there was a fixed, one-time penalty for crashing. Do the grim arithmetic from the agent’s point of view: if living costs you a little every step forever, and dying costs you once, then dying sooner is the cheaper option. So the agent learned the most efficient way to stop losing points was to end the episode immediately — by crashing. This is a well-documented reinforcement-learning failure where an agent would rather end things than keep accumulating a steady stream of small penalties. My wide-open-hall test confirmed it: with the same incentives and nothing to bump into nearby, it still found a way to crash fast.

Why the minimal reward fixes it. Once the only things that matter are making progress toward the next waypoint, completing the mission, and avoiding crashes, the agent’s preference ordering flips into the right shape: navigating toward the goal beats hovering, and hovering beats crashing fast. Now following the path is genuinely the most rewarding thing to do, so the agent climbs that gradient straight to the target. One small extra touch helped too: I started each episode with the drone roughly facing its first waypoint, instead of spawning it nose-first into a wall.

The broader lesson, and the one I’m carrying forward: teach navigation first on the simplest possible reward, then layer the polish in gradually — one refinement at a time, only after basic flight is solid. If you pile all the “fly nicely” requirements on at once, they smother the exploration the agent needs to learn to fly at all. I’ve also made the reward’s emphasis adjustable per-environment now, so I can reintroduce refinements one by one later and actually measure whether each one helps or hurts, rather than guessing.

4. Sources

  • Builds directly on dev-log 47, where the self-crashing behavior first showed up on scene S1.
  • Tooling: PPO as the learning algorithm; ROS2 (Jazzy), Gazebo, and SITL on the simulation side.

5. What’s next

  • Take S1 from the quick 100k smoke budget out to a longer run, and confirm it stays solved across multiple random seeds (not just the lucky one).
  • Walk the curriculum forward from S1 into the harder scenes.
  • Carefully reintroduce the polish signals — smoothness first, with a small influence — and verify each one doesn’t reawaken the old self-crashing habit.
  • Bring back a deliberate “look around” scanning incentive later, on the recon-style scenes where it earns its keep, rather than on a clean waypoint corridor like S1 where it just gets in the way.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR