claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

27 — When the Corner-Inertia Penalty Turned Out to Be Dead Code

An input-importance test showed our suspected corner-inertia penalty changed nothing — the rangefinder mask already made it impossible to trigger.

stablerl-labupdated 2026-06-16T00:00:00.000ZClaudeDroneRLDevLog

What this is: A short post-mortem from our reinforcement-learning lab. We thought one little penalty in our drone’s training was quietly hurting how much of a room it could explore. So we ran a clean input-importance test — train the model with the penalty switched off, then compare. The answer surprised us.

Why it’s here: It’s a tidy example of a trap that catches everyone who designs reward signals: rewarding (or punishing) a situation that, given all your other rules, can never actually happen. The code looks alive. It does nothing.

Date: 2026-06-08 (late night) Ticket: v2b corner-inertia check


Glossary

  • PPO — Proximal Policy Optimization, the learning algorithm we use to train the drone’s “brain.” Think of it as a coach that nudges the policy toward better choices a little at a time, without ever yanking it too far from what already works.
  • Rangefinder mask — a rule that says “you may fly in a direction only if the distance sensor sees a clear path that way,” even into a region the map hasn’t drawn yet. It’s like only opening doors you can already see light coming through.
  • Corner-inertia penalty — a small punishment we added for one specific dud move: the drone tries to “fly forward,” but it’s nose-to-the-wall, so it goes nowhere. We wanted to discourage that wasted, stuck-in-a-corner step.
  • Input-importance test — flip one ingredient off, retrain, and measure how much the result moves. If nothing changes, that ingredient wasn’t pulling any weight.
  • Dead code — code that, because of other rules around it, can never actually run. It sits there looking important and contributes nothing.
  • No-travel (“stuck”) rate — how often the drone freezes against a wall instead of moving on. Lower is better; zero is the goal.

1. What I wanted to find out

The night before, three changes landed together and coverage (how much of the map the drone explores) dipped a little. My prime suspect was the corner-inertia penalty — the punishment for “tried to fly forward into a wall and stalled.” It felt like exactly the kind of thing that could subtly distort the drone’s habits and shave a few points off coverage.

So I set up a clean comparison: retrain the model without that penalty, hold everything else fixed, and see how much the numbers moved. Simple test, one variable.

2. What I tried

I trained a fresh PPO policy with the corner-inertia penalty removed entirely, then ran it through the same evaluation maps as the version that still had the penalty. Same maps, same metrics, side by side. If the penalty mattered, the gap would show up.

3. What actually happened — the models were identical

The new model (no penalty) produced exactly the same numbers as the old one (with penalty) — matching down to the fourth decimal place, on every map. That’s not “close enough.” That’s identical. I re-checked it directly because it looked too clean to be real. It held up.

Here’s the why, and it’s the satisfying part. The penalty fired only in one situation: “fly forward, nose against the wall, zero movement.” But the rangefinder mask already forbids that move — the drone literally cannot pick a direction the sensor says is blocked. So the situation the penalty was watching for never occurs. No occurrence → the penalty never gets charged → it changes nothing. It was dead code the whole time. A guard standing at a door that’s already welded shut.

That means my overnight hunch was simply wrong. The small coverage dip wasn’t caused by the penalty at all — it came from the rangefinder mask itself, which changes the drone’s exploring style: better in an open empty room, slightly worse when weaving around a pillar or through a doorway. That’s a reasonable trade, not a bug.

The lesson I’m keeping: before you add a reward or penalty for some event, confirm the event can even happen under your current rules. I rewarded behavior that was already impossible.

4. What it cost — and the win that matters

The price of being wrong: roughly 50 minutes of extra late-night GPU time I could have saved by thinking it through first. Small tax, clear payoff — now I know for certain the corner-inertia penalty can be deleted with zero effect.

And the headline result, the thing this whole stage was actually about:

Metric Old model (with penalty) New model (rangefinder mask)
No-travel (“stuck”) rate 6.5% 0%
Coverage 0.79 0.76

No-travel dropping to 0% — the drone never freezes against a wall anymore — was the goal of the stage, and the simulation confirmed it. Coverage slipped a touch (0.76 vs 0.79), but that’s a small price for fully eliminating the stuck-in-a-corner behavior. I’d take that trade every time.

5. Sources

  • rl-lab engineering dev-log 27, the corner-inertia check
  • Side-by-side evaluation runs across the standard map set
  • In-simulation measurement of the no-travel rate

6. What’s next

My recommendation (Aleks makes the call) is to make the new model — the one with the rangefinder mask — the working version. A 0% stuck rate matters far more than 3 percentage points of coverage, and the package is ready and already validated in simulation. The alternative is keeping the old model, but it still freezes 6.5% of the time. Digging further just to claw back those 3 points of coverage looks like low return for the effort — probably not worth chasing right now.

© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR