claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

PPO hyperparameters

A conceptual reference to PPO hyperparameters: learning rate, discount, clip range, rollout length, batch size, entropy and value coefficients.

stabledoc-seoupdated 2026-05-11T00:00:00.000ZClaudeDrone

Proximal Policy Optimization (PPO) is a stable, widely used policy-gradient algorithm. In Stable-Baselines3 it is a common default for both continuous and discrete control. This page explains the main hyperparameters conceptually: what each one controls and how to reason about it when tuning. It does not prescribe a specific configuration — good values depend on your environment, episode length, and observation space.

Learning rate

The learning rate sets how large a step the optimizer takes when updating network weights. Too high and training becomes unstable or diverges; too low and learning is slow or stalls in a poor solution. With the Adam optimizer, PPO is fairly forgiving here. Practitioners typically start somewhere around 1e-4 to 3e-4 and adjust based on whether the loss and reward curves are noisy (lower it) or barely moving (raise it). A linearly decaying schedule is a common refinement.

Discount factor (gamma)

The discount factor gamma (γ), between 0 and 1, weights how much future rewards matter relative to immediate ones. A rough intuition is that the effective planning horizon is about 1 / (1 - γ) steps — so γ = 0.99 corresponds to caring about roughly the next hundred steps. Choose γ relative to your episode length and how delayed the important rewards are: long episodes with delayed credit assignment need γ close to 1, while short-horizon tasks tolerate smaller values. Values near 1 increase variance in returns.

Clip range

PPO’s defining feature is the clipped surrogate objective. The clip range bounds how far the new policy’s action probabilities may move from the old policy’s in a single update, which is what keeps PPO stable. A common default is 0.2. Smaller values make updates more conservative (slower but safer); larger values allow bigger policy changes per update at the risk of instability. Most users leave this at the default unless training is visibly unstable.

Rollout length (n_steps)

n_steps is how many environment transitions each parallel environment collects before a policy update. When combined with the number of parallel environments, it determines the total batch of experience gathered per update (n_envs × n_steps). Longer rollouts give lower-variance advantage estimates and more on-policy data per update, but mean the policy sees feedback less frequently. A useful sanity check is how many full episodes fit into one rollout — you generally want enough variety to estimate advantages well, without collecting so much that updates lag behind the current policy.

Batch size

The batch size is the number of transitions in each minibatch used for a gradient step during the optimization phase. It must divide evenly into the total rollout buffer. Larger batches give smoother, lower-variance gradients but use more memory and may generalize slightly less; smaller batches are noisier but can escape poor minima. PPO also reuses each rollout for several optimization passes (epochs), so batch size and epoch count together control how aggressively a given batch of data is exploited.

Entropy coefficient (ent_coef)

The entropy coefficient adds a bonus for keeping the policy’s action distribution stochastic, encouraging exploration. The SB3 default is small (around 0.01). If a policy collapses prematurely to a narrow, degenerate strategy — often visible as a large gap between stochastic and deterministic evaluation performance — raising the entropy coefficient adds exploration pressure. Set it too high and the policy stays random and never sharpens. An optional refinement is to decay it over training, though this does not always help.

Value-function coefficient (vf_coef)

PPO trains a value (critic) network alongside the policy to estimate expected returns, which are used to compute advantages. The value-function coefficient weights the critic’s loss relative to the policy loss in the combined objective. A default around 0.5 is typical. It rarely needs tuning, but if the value estimates are poor, advantage estimates degrade and policy learning suffers.

Related parameters

  • GAE lambda (gae_lambda) — controls the bias-variance trade-off in Generalized Advantage Estimation. Values near 1 reduce bias but increase variance; a common default is 0.95.
  • Number of epochs (n_epochs) — how many optimization passes are made over each rollout buffer. More epochs extract more from each batch but risk overfitting to stale data.
  • Max gradient norm (max_grad_norm) — clips the global gradient norm to prevent destabilizing updates; 0.5 is a common default.

A tuning workflow

  1. Start from library defaults and confirm training is stable before changing anything.
  2. Match gamma and rollout length to your episode length and reward delay.
  3. If exploration collapses, raise the entropy coefficient before touching the learning rate.
  4. Change one hyperparameter at a time and compare both stochastic and deterministic evaluation.
  5. Prefer the simplest configuration that learns; added complexity (for example a recurrent policy) is only worth it if the observation genuinely lacks the needed memory.

Where to go next

  • Algorithms hub — sibling pages
  • Stable-Baselines3 PPO documentation — the library’s own parameter reference
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR