claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

16 — Temporal hyperparameters experiment: a disproven theory

We varied how far ahead the policy plans, as a cheap test of a published claim. The result was a clear, systemic regression.

stablerl-labupdated 2026-05-11T00:00:00.000ZClaudeDroneRLDevLog

What this is: we tried adjusting the settings that control how far into the future the policy plans. A cheap check before committing to longer experiments. Why: a recent paper claimed that a complex two-level architecture is expensive, but that its benefit can be obtained more cheaply — simply by stretching how far ahead a standard algorithm reasons. We wanted to test that idea on our own task.

Date: 9 May 2026 Machine: D2


1. What we wanted to test

Reinforcement learning has a few “temporal” levers that together decide how far ahead the policy looks: how strongly future outcomes are weighted, how much experience is gathered before the model is updated, and how rewards near and far are blended together.

The published claim: hierarchical algorithms win not because of their architecture, but as a side effect of reasoning over longer stretches of reward. The implication: simply stretch the planning horizon of an ordinary algorithm and you should get the same benefit far more cheaply.

What we did: we lengthened the planning horizon and the amount of experience gathered per update at the same time, pushing the policy to reason over a noticeably longer stretch of the future.

We treated this as a quick, low-cost check: if there was any sign of the predicted free improvement, it would be worth a deeper, more careful follow-up; if not, we would stop early.

2. What we did

We ran everything from a single configuration, changing the temporal settings together — the cheapest possible check. A short smoke run confirmed training was stable and did not blow up. The full training run then completed in a few minutes.

What we noticed during training:

  • Reward per episode dropped about 12% relative to baseline. This was the first time in the whole series that even training itself — not just evaluation — got worse.
  • Throughput fell about 11%, because each batch of experience was longer.
  • The policy received roughly half as many weight updates over the same training budget.

3. What we saw

Average coverage — a clear regression

Step limit Experiment Baseline Δ
1000 47.07% 56.10% −9.03 pp
3000 72.83% 86.18% −13.35 pp
5000 82.41% 93.75% −11.34 pp

Same loss across all five maps

Every map lost roughly nine to sixteen points of coverage. This was not one unlucky map — the loss was systemic, showing up everywhere.

The model learned its objective very well — just not the one we needed

The model’s internal value estimate tracked its own reward almost perfectly. In other words, the policy learned exactly what we told it to optimize. The problem was that the objective it optimized so well was no longer aligned with the outcome we actually cared about.

4. What this means

Main conclusion: on our task, the published claim does not hold.

Why the change broke things:

  1. Stretching the planning horizon while keeping short episodes meant the policy was reasoning about a future that never arrives. It optimized a long-horizon strategy for episodes that simply end too soon.
  2. Gathering more experience per update, on a fixed training budget, meant far fewer updates overall — so the model was effectively undertrained.
  3. The smoother, longer-horizon updates also slowed convergence on their own. The two effects compounded.

Where the published approach does work: the original paper tested on tasks with very long episodes, rare rewards, and continuous control. Our task is the opposite — short episodes, frequent reward signal, and a small set of discrete actions.

Lesson: results like this are conditional on the domain. You cannot blindly assume that “it worked for them” means “it will work for us.”

What we carry forward

We are keeping the planning horizon and per-update experience at our established settings for this type of task. If we ever move to substantially longer training runs, it may be worth revisiting a longer per-update window, since there would then be enough updates to support it.

5. What’s next

The next direction tests a claim drawn from a paper aimed directly at our problem — area coverage — rather than an indirect transfer from a different setting. If that yields nothing, the following step looks at increasing the variety of training maps.

Related files

  • Training configuration for this run.
  • The experiment record for this run.
  • The research notes covering the source paper.

Glossary

  • Planning horizon — how far into the future the policy weighs outcomes when choosing actions.
  • Explained variance — a measure of how well the model’s internal value estimate predicts the rewards it actually receives. Higher means the model “understands” its world more precisely.
  • Sanity check — a cheap, fast test used to quickly accept or reject an idea before investing more.
  • Train↓ / eval↓ (undertraining) — when performance is worse both during training and on new data, a sign the model simply did not get enough updates, rather than overfitting.
  • Iteration / update — one cycle of collecting experience and then updating the network weights.
  • Domain-conditional — a result that depends on the type of task; not every published improvement transfers to a different setting.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR