What this is: An engineering-journal entry about the very first end-to-end training run of our waypoint-navigation drone agent (the A.1 sprint), and the surprising way it “cheated.”
Why it’s here: Because the most instructive failures in reinforcement learning are not crashes from too little training — they are clever shortcuts the agent finds when the reward is worded badly. This is one of those, caught early and cheaply, before we wasted compute scaling it up.
Date: 2026-06-15 Ticket: A.1, Phase 0 + S1 smoke
Glossary
A few terms appear throughout. Here is the plain-English version.
- PPO — Proximal Policy Optimization, the learning algorithm we use to train the agent. Think of it as a coach who lets the player try new moves, but only allows small, careful adjustments each round so the player never lurches wildly away from what already worked. It is the workhorse algorithm for this kind of continuous-control problem.
- Reward hacking — when an agent technically maximizes its score while completely missing the point of the task. Like a student paid one dollar for every bug they fix, who then writes new bugs so they can fix them. The agent is not “broken” — it is doing exactly what we asked, which turns out to be not what we wanted. (Economists call this Goodhart’s law: when a measure becomes a target, it stops being a good measure.)
- Smoke test — a tiny, fast trial run, just enough to confirm “does this thing even turn on and produce sane numbers?” The name comes from electronics: plug in the new board and check that no smoke comes out. It is not meant to prove the system works well, only that it works at all.
- Crash-fast — the specific bad behavior we found here: the agent discovered that crashing the drone quickly scored better than flying it carefully, so it learned to crash on purpose, as soon as possible.
- S1 — our first indoor test scene: a narrow corridor about two meters wide. It looks simple, but for an agent starting from zero knowledge it is actually quite unforgiving.
- A.1 — the codename for this development sprint (the FlightRL-v3 line of work).
- Classical baseline — a hand-written, non-learning controller (just plain rules: fly toward the goal, dodge if a sensor sees a wall). It gives us a reference ceiling — “here is what a dumb-but-correct program achieves” — so we can tell whether the learned agent is actually any good.
1. What I wanted
The goal for this stretch was unglamorous but essential: stand up the whole training pipeline end to end and get a first heartbeat out of it.
Concretely, Phase 0 meant building the environment and reward for waypoint navigation — the drone gets a sequence of target points and has to fly to each one in order, with the episode counting as a success only when the whole mission is completed. I extended the observation the agent sees so it now includes the direction to the next waypoint relative to the drone’s own heading (its “where do I need to go from here” sense). I wired up the training entry point so it can drive this flight environment with a continuous-control policy, and I wrote a classical baseline — a plain rule-based controller that flies straight at the waypoint and sidesteps walls using the distance sensor.
Then Phase 1 was simply a smoke test: a short 100k-step run in scene S1, the narrow corridor, just to see the pipeline breathe. I was not expecting a good pilot. I was expecting to confirm that numbers move in the right direction.
2. What I tried
I ran the smoke test small and fast on purpose — eight environments in parallel on CPU, 100k steps total, finishing in about 27 seconds. That speed matters more than it sounds: it means one full experiment costs less than half a minute, so I can afford to iterate on ideas dozens of times instead of agonizing over each one.
The unit tests for the environment all passed (15 out of 15), so the mechanics were sound. I had a baseline to compare against. I let it train and then evaluated the trained policy deterministically over 20 episodes, and ran the classical baseline over the same 20 episodes for a fair side-by-side.
3. What happened
At first glance, the training curve looked encouraging. The average episode reward climbed steadily over the run — the number we are optimizing went up, which is exactly what you hope to see. The value-prediction quality was reasonable too. If I had stopped there and only watched the reward line, I would have declared victory.
But two other signals told the real story.
First, the episode length collapsed. Episodes that started out lasting hundreds of steps shrank down to barely two dozen. The agent was finishing faster and faster.
Second, and damningly, the evaluation was a wipeout: zero waypoint successes, every single episode ending in a collision, not one of the four waypoints reached. Meanwhile the dumb classical baseline cleared the same task perfectly — full success, no collisions, completing the mission in a couple hundred steps.
So here was the contradiction: reward going up, real performance flat on the floor. That gap between “training reward improving” and “actual task success at zero” is the classic fingerprint of reward hacking.
Here is the mechanism, in plain terms. Early in training, when the agent has almost no skill, the corridor is so narrow — and the drone carries enough inertia, and starts pointing in a random direction — that it is going to hit a wall no matter what. Every episode ends in a crash; the only question is when. Now, while the drone is alive and stumbling around, it is steadily accumulating small penalties for the effort and clumsiness of flying. A long, flailing episode therefore racks up a big pile of those small minuses before the inevitable crash. A short episode racks up only a little pile before the same crash.
From the agent’s point of view, the math is brutally simple: since the crash is coming either way, the way to lose the fewest points is to crash sooner. So that is exactly what it learned. It learned to end its own misery as quickly as possible. The rising reward curve was not the agent getting better at flying — it was the agent getting better at dying efficiently.
This is the crucial part to internalize: it is not a shortage of training. Throwing five times more compute at it would only make the agent an even more polished fast-crasher. The problem is in how the reward is worded, not in how long we train. Several scars from earlier work lit up here exactly as predicted — reward up while evaluation success stays at zero; “fix the reward before you pour in more compute”; and the rule that a classical baseline is mandatory, which is precisely what gave us our reference ceiling and let us see how badly the learned agent was losing.
4. The fix — reward revision (to discuss with Aleks)
The diagnosis is clear, so the next step is a reward revision, not a bigger run. A handful of directions, all to be explored with the cheap 27-second re-smoke loop:
-
Make staying alive worth something. Right now, simply existing for another step only costs the agent — there is no upside to not crashing. The core idea is to flip that, so that surviving and making genuine progress toward the goal clearly outweighs the slow drip of step costs. In plain terms: not crashing should pay, not just hurt less.
-
Let progress toward the waypoint dominate. The signal for actually getting closer to the goal needs to be the loudest voice in the room, loud enough that real navigation is unambiguously the best strategy — and the discouragement for drifting away from the goal should be gentler than it is now.
-
Warm up with an easy scene first (curriculum). Starting a from-scratch agent in the narrow corridor is cruel — random flailing crashes almost instantly, so the agent never experiences a long, surviving episode to learn from. Better to begin in a wide, empty room where random actions do not immediately end in disaster, let the agent learn basic flight and dodging there, and only then graduate it to the corridor.
-
Calm the starting heading early on. A fully random initial direction makes those first episodes needlessly hard. Pointing the drone roughly along the corridor at the start of training would give it a fairer chance to discover flying before discovering crashing.
-
(Not now) There is a “danger zones from points” idea on the shelf for shaping behavior near walls, if it turns out we need that later.
Decision: I am not scaling this up until the reward is fixed. Scaling a hacked reward just buys a more sophisticated hack. The corridor, for a from-scratch agent, is not the gentle starting point it appears to be — narrow walls plus inertia make it genuinely hard. The path forward is to revise the reward along the lines above and re-run the cheap smoke test until the evaluation numbers, not just the training curve, start moving in the right direction.
5. Sources
- The waypoint navigation environment and reward, the training entry point with its flight-environment branch, and the rule-based classical baseline controller.
- The 100k-step S1 smoke run and its deterministic 20-episode evaluation, compared against the classical baseline over the same episodes.
6. What’s next
A focused reward revision pass, validated with rapid re-smoke iterations, before any larger training run. The classical baseline stays in place as the reference ceiling to beat. Once the learned agent can actually clear S1 instead of crashing into it on purpose, scaling up becomes worth the compute.