Staleness and off-policy drift in asynchronous LLM RL — related-works note

Reviewed 2026-09-17. Staleness is the version lag between the policy that sampled a trajectory and the policy whose gradient that trajectory updates. In asynchronous post-training it is the default condition, not a corner case. The 2025–2026 recipes treat it as a controlled quantity: bounded by admission control or version-drop rules, corrected at the loss with ratios taken against the inference engine's own log-probabilities, and monitored through log-probability-difference and importance-ratio statistics. No source reviewed here calls the problem solved. Two confounds must be separated before any result is attributed to staleness: training-versus-inference numerical mismatch, which is present at zero lag, and multi-gradient-step mini-batching, which creates staleness with no asynchrony at all.

This is a source-pinned reading note, not an execution study. No training runs were launched. Local pins: the SkyRL checkout at ../SkyRL (commit 152ebec0, 2026-06-25), the GLM-5 report text extracted from its PDF, and the RL-Infra-Bench blueprint. Per-source records, with which numbers were re-read on the primary page and which come from a delegated sweep, are in four companion notes: systems, estimators, mismatch and the 2026 recency sweep.

1. Definitions and the decomposition

Let μ be the sampling distribution the inference engine actually realized, π_old the trainer's recomputation of the same weights, and π_θ the policy being updated. The token-level importance ratio an on-policy surrogate needs factors as

π_θ(x_t) / μ_old(x_t)  =  [ π_old(x_t) / μ_old(x_t) ]  ×  [ π_θ(x_t) / π_old(x_t) ]
                            training–inference mismatch      policy staleness

The decomposition appears in the Qwen team's formulation paper, in SkyRL's off-policy-correction documentation (crediting verl's rollout-correction math note), and as r = r_d·r_s in the May 2026 "missing old logits" paper. The first factor is nonzero in a perfectly synchronous run because inference kernels, precision, parallelism and MoE routing differ from the trainer's. The second is what asynchrony, in-flight updates, partial rollouts and mini-batching add. The Qwen paper's first-order argument is that the token-level surrogate tracks the true sequence-level objective only while both factors stay small. Qwen, arXiv 2512.01374 · SkyRL docs · Guan et al., arXiv 2605.12070

Staleness is counted in policy versions, and the accounting differs by system. SkyRL counts the training step that consumes a group minus the step at which it was scheduled. AReaL, ROLL Flash and StaleFlow count from the version that started generation. GLM-5 records the ordered list of versions that touched a response and measures from the oldest. ESTR separates intra-trajectory staleness (versions spanned inside one rollout under in-flight updates) from inter-trajectory lag (last token's version to the target policy). Applied Compute splits mean staleness into a pre-queue part accrued during generation and an in-queue part accrued while waiting for the trainer. Numbers from different papers are therefore not comparable without restating the definition. ESTR · Applied Compute

Under in-flight updates the behavior policy is not a checkpoint. GLM-5 states that tracking exact behavior probabilities across the versions a single response spans would require keeping every intermediate checkpoint, and instead uses the rollout engine's logged log-probabilities as the behavior proxy. The missing-old-logits paper measures what recovering π_old exactly would cost (snapshots: 40–76 GB and 25–150 s per step on 4B and 30B models) and proposes an exponentially weighted parameter average as the proxy. GLM-5 §4.1.2, arXiv 2602.15763

2. Where staleness comes from

Source Mechanism Pinned example
Decoupled generation and training Engines generate continuously; the trainer consumes whatever finished; weights are pushed back periodically GLM-5 pushes every K gradient updates and resets the optimizer after each push; Laguna broadcasts every 2 optimizer steps; ECHO-2 publishes every S−1 steps
In-flight weight update An unfinished sequence continues under new weights PipelineRL keeps the stale KV cache; AReaL recomputes it; INTELLECT-3 lets one trajectory span several policies
Partial rollouts Long trajectories are paused at a budget and resumed next iteration Kimi k1.5/K2/K3 pause once a fraction λ completes; APRIL resumes over-provisioned rollouts across up to five versions with no correction
Long-tail trajectories A group scheduled at step i finishes after several steps SkyRL's capacity rule bounds aggregate lag, not per-group lag; violators are accepted with a warning
Mini-batching and batch reuse Several gradient steps per sampled batch make later steps off-policy with no asynchrony Qwen varies N ∈ {1,2,4,8} mini-batches; μ-GRPO and PNPO reuse batches for hundreds of updates
Stale KV cache Prefix blocks computed under an older policy are reused after a sync SkyRL clear_kv_cache_on_weight_sync defaults to false in code; Laguna resets the cache on every broadcast
Text round-trips Re-tokenizing engine text breaks token-level correspondence and masquerades as off-policyness GLM-5 token-in-token-out gateway; Mercor rewrote its harness to use /completions

3. How systems bound staleness

Admission control. AReaL rejects new generation requests when ⌊(N_r−1)/B⌋ > i+η, with N_r trajectories generated so far, B the batch size and i the current version; η = 8 for math and 4 for code. SkyRL reimplements this as a capacity rule, accepted + running < (max_staleness_steps + current_step) × mini_batch_size, with a buffer of mini_batch_size × (S+1) finished groups and the constraint mini_batch_size ≤ workers ≤ mini_batch_size × (S+1); default S = 4. ROLL Flash caps its sample buffer at (1+α)×batch so no generated sample can violate the bound. StaleFlow reserves a version tag at admission and consumes a trajectory only if V_traj + η ≥ V_buf. All four are aggregate or per-trajectory bounds on versions, not on wall-clock. AReaL · fully_async_trainer.py:104–176 · ROLL Flash · StaleFlow

Drop rules at training time. GLM-5 discards a response when the current version minus its oldest version exceeds τ (unreported). INTELLECT-3 discards rollouts generated by more than max_off_policy_steps = 8 policies. ECHO-2 discards below v_t − S. verl's V1 sampler drops, or optionally waits on, prompt groups at max_off_policy_threshold = 8 versions; NeMo-RL evicts trajectories older than max_trajectory_age_steps (default 1); PipelineRL's source counts samples beyond max_lag rather than dropping them, although the March 2026 Hugging Face survey classifies it as version rejection. SkyRL's documentation prefers proactive admission to over-generating and dropping, on the grounds of wasted compute.

In-flight handling is where implementations diverge most. PipelineRL continues in-progress sequences with the stale KV cache and reports only slightly higher divergence than recomputing it. AReaL discards and recomputes the cache. The June SkyRL checkout freezes requests with vLLM's keep-mode pause and resumes them after the sync; its own tutorial and the Hugging Face survey still describe an abort-and-refeed loop, so the behavior changed between March and June 2026. Laguna blocks any in-flight rollout step on an update so no step straddles a refresh, and resets the KV cache. ECHO-2 workers simply switch version when a snapshot finishes installing. The Prime Intellect 1T run drops by max_off_policy_steps and salts the KV cache by version. PipelineRL · inference_engine_client.py:381–395 · Laguna, arXiv 2605.27605

What sets the mean staleness. Applied Compute derives closed forms in five parameters: trainer utilization ρ, batch size B in rollouts, sampling concurrency C, queue factor q and a response-length tailness multiplier. Rollout-bound: staleness ≈ C·M_tail/B + ρ. Train-bound: C·M_tail/(ρB) + (2q+ρ−1)/(2ρ). Predictions on real Qwen3-8B GRPO runs were within 0.06–0.27 versions of measurement; setting q = 1 cut staleness 35% at unchanged step time in one configuration. This is the same dimensionless-ratio framing the RL-Infra-Bench blueprint uses, now with a validated formula. Applied Compute, 2026-07-03

Production ceilings. Mercor's SkyRL runs use S = 3 with mini-batch 16 and 16 samples per prompt, a 1,024-trajectory ceiling, run at 550 (35B) and 300 (397B) concurrent rollouts. Laguna caps at 10 optimizer steps and never reaches it. GLM-5 runs more than 1k concurrent rollouts with τ and K unreported. Marin's 67B runs recommend a staleness allowance of 4 updates. Mercor guide, 2026-09-01

4. Loss-level corrections

All of these take the ratio against the rollout engine's logged log-probabilities. They differ in what they do with it.

Correction Rule Where pinned
Truncated importance sampling, token multiply the loss by min(π/μ, C) per token; SkyRL default C = 2.0; verl default threshold 2.0; PipelineRL clamps at 5; LlamaRL's AIPO uses C in [2, 10] SkyRL tis_ratio_type="token"; Yao et al.
Truncated importance sampling, sequence same with the product over tokens; SkyRL default C = 5.0; Qwen uses cap 5; VCPO uses 8 SkyRL tis_ratio_type="sequence"; Liu et al.
Sequence mask, geometric or product zero a sequence when the geometric mean of ratios leaves [0.99, 1.01] or the product leaves [0.5, 2.0] SkyRL sequence_mask_metric; verl MIS mode; DeepSeek-V3.2 masks negative-advantage sequences whose mean log-ratio exceeds δ
Outlier-token sequence mask zero a sequence if any token ratio is outside [1e-4, 100]; verl vetoes below 1e-4 SkyRL outlier_token_is_threshold_*
Token mask (IcePop family) zero individual tokens with ratio outside [1/β, β]; Ring-1T and INTELLECT-3 use [0.5, 5]; GLM-5 synchronous RL uses β = 2 GLM-5 §3.2; SkyRL token_mask_is_threshold_*
Direct double-sided IS (GLM-5 agentic loss) r_t = π_θ/π_rollout with no π_old; gradient f(r_t) Â log π_θ, f zero outside [1−ε_l, 1+ε_h] GLM-5 §4.1.2 eq. 3–5; SkyRL policy_loss_type="rollout_is"
CISPO clip the weight, not the update, and keep the gradient through log π; Laguna's asymmetric clip gives an effective range [0, 5] MiniMax-M1; ScaleRL; Laguna
DPPO binary mask on the probability difference π_θ − μ (total variation) or a binary KL, thresholds 0.2 or 0.05, rollout log-probs as behavior policy arXiv 2602.04879; SkyRL policy_loss_type="dppo"; Mercor
Second-moment trust region (M2PO) no ε-clip; mask the highest (log r)² tokens until the batch mean of (log r)² is below 0.04 arXiv 2510.01161
Entropy-scaled trust region (ESTR) keep a token iff δ_t²/(H_t+ε) ≤ τ with H_t the behavior entropy arXiv 2607.22186
Three-policy objectives PPO-clip π_θ/π_prox times π_prox/π_behav (AReaL decoupled PPO); verl decoupled mode; A-3PO interpolates π_prox instead of recomputing it AReaL; verl docs
Optimizer-side scale the learning rate by √(ESS_off/ESS_on) (VCPO); project or skip updates by gradient cosine (GAC); S·η_max ≈ 1.6e-6 as an empirical law VCPO; arXiv 2607.01083

SkyRL applies its masks multiplicatively in a fixed order (outlier sequence mask, token mask, sequence mask) and logs is_ratio_{mean,std,max,min}, clip fractions and rollout_train_logprobs_abs_diff_*. Every correction is silently disabled when no rollout log-probabilities are attached to the batch. off_policy_correction_utils.py:257–312

5. How much staleness is tolerable: reported evidence

Study Setup Finding Boundary
Asynchronous RLHF (Oct 2024) Pythia 410M–2.8B; N ∈ {1..64} mini-batch updates per generation online DPO still learns at N=64; PPO and RLOO degrade sharply; larger policies more robust production setting is one-step lag only
AReaL (May 2025) 1.5B, AIME24, η ∈ {0,1,2,4,8,∞} naive PPO 42.0 → 23.3 at η=4; decoupled PPO 42.0 / 42.1 / 41.8 / 42.2 / 41.0 / 36.9; about 2.1× throughput at η=4 one model, one benchmark
ScaleRL (Oct 2025) 8B dense and 17B×16 MoE; PipelineRL-k, k ∈ {1,8} k=8 chosen; same asymptote as PPO-off-policy-8 with better compute efficiency; FP32 head raises asymptote 0.52 → 0.61 two k values; the precision result is a mismatch effect
ROLL Flash (Oct 2025) Qwen3-8B-Base; α sweep α=2 optimal for most settings; decoupled PPO, TIS, CISPO, TOPR and vanilla GRPO comparable at α=2 and α=8 curves only
M2PO (Oct 2025) 1.7B–32B; data delayed by k ∈ {0..256} updates Qwen2.5-Math-7B: GRPO 49.3 on-policy, 45.7 at k=256; GSPO 44.7; M2PO 48.8; decoupled PPO, TOPR, CISPO break down early at k=256 staleness simulated, not a live pipeline
Qwen "Stabilizing RL" (Dec 2025) Qwen3-30B-A3B, FP8 inference, N ∈ {1,2,4,8} mini-batches N=2 needs routing replay plus clipping; N ≥ 4 needs full R3 plus clipping or collapses staleness from mini-batching; math only
ECHO-2 (Feb 2026) Qwen3-8B; S ∈ {3,4,6,11} S ≤ 6 within about 5% of synchronous; S = 11 diverges consumer-GPU, wide-area regime
VCPO (Feb 2026) Qwen2.5-7B; PipelineRL-k with k = 10–12, stress to 128 MATH-500 72.0 sync vs 71.6 at k=10 with 1.45× fewer GPU-hours; TIS, MIS, M2PO collapse at k=12 on GSM8K ≤7B
StaleFlow (Jan/Aug 2026) Qwen family to 30B-A3B; η ∈ {0,1,3}, 10 η=3 matches synchronous convergence; η=10 or no control collapses single η in the convergence check
Huang (May 2026) toy bit-sequence MDP, horizon 1024 every method collapses at K=12; sequence-level TIS matches sync from batch 32 and exceeds it at 128 while token-level keeps collapsing synthetic
Staleness–LR law (Jul 2026) Llama-3.2-1B/3B, S ∈ {8,16,32} η_max ∝ 1/S with S·η_max ≈ 1.6e-6; collapse time t·η ≈ 3.2e-5 regardless of S tiny models; local analysis
ESTR (Jul 2026) Qwen3-30B, Qwen2.5-7B; Δ_intra to 9, Δ_inter to 30 IcePop and KPop collapse within a few hundred steps on multi-turn GSM8K; ESTR 95.69; 2.6× throughput τ not captured; no TIS/MIS baselines
Marin standup (Sep 2026) 67B; staleness 0/1/2/4 updates 1: −30% wall-clock, +0.3 quality (5 seeds); 4: −55%, −2.6 (3 seeds), judged non-inferior informal; acceptable range "not established" per the companion issue
Mercor/SkyRL (Sep 2026) 35B and 397B; S=3 DPPO within noise of the GLM-5 loss at epoch 1 but shorter, more numerous turns (21→32 turns; 834→588 tokens per turn); mean log-prob diff under 0.03 treated as healthy one run per size

The pattern: tolerated staleness moved from about one step in 2024 to 3–8 versions in 2025–2026 production recipes, and to 50–256 updates in method papers on small models, but only after ratios were taken against engine log-probabilities, masks were added, and MoE routing was replayed. Quote a tolerance only together with its loss, its mismatch controls and its staleness definition.

6. The mismatch confound

Mismatch is measurable at zero staleness and is often the larger factor. Measured magnitudes: KL(engine‖trainer) roughly 1e-4–1e-3 in bf16 on dense models, 1e-2 in FP8, up to 1e-1 for MoE; a 30B-A3B MoE flips at least one expert on 94% of tokens between SGLang and Megatron. Fixes and costs:

Two August 2026 studies test whether removing mismatch helps under staleness. Wang's GDN study reaches an exactly zero gap at lag 0 with batch-invariant kernels, then sweeps off-policy windows of 4, 12 and 32: at window 12 on MATH the parity arm peaks near 0.7 train reward against 0.6 for vLLM native, but at window 4 the arms interleave, on Search-R1 they are indistinguishable, and at window 32 both degenerate. The author's conclusion is to use parity as a debugging control, not a production default. Panda's MoE study gets K3 to zero on Qwen3.6-35B-A3B and raises held-out solve from 63.9% to 77.4% synchronously, but an asynchronous run with stale weights and cache and no mitigation falls to 59.4%, recovering to 76.6% with router recall plus CISPO. Both confirm the ordering in §1: fix numerics first so that any residual ratio is attributable to lag, then bound and correct the lag. Wang, Aug 2026 · Panda, 2026-08-17 · ../torchtitan-batch-invarient-GDN

7. Framework knobs

Framework Async mode Staleness knob and default Correction options In-flight handling
SkyRL (local, June 2026) trainer.fully_async max_staleness_steps = 4; num_parallel_generation_workers = 768; clear_kv_cache_on_weight_sync false in code, "true" in the config page none by default; off_policy_correction.*, rollout_is, dppo, cispo, gspo, sapo keep-mode pause; requests frozen and resumed
verl V1 trainer (sync, colocate_async, separate_async); fully-async recipe; one-step-off V1 trainer.v1.sampler.max_off_policy_threshold = 8 with drop or wait; fully-async async_training.staleness_threshold = 0 (a fraction; keep below 1) algorithm.rollout_correction: TIS or rejection × token, sequence, geometric; threshold 2.0; bypass_mode; IcePop preset V1 keeps tokens and log-probs, rebuilds KV, resumes on the new version
AReaL native async rollout.max_head_offpolicyness = 0 (docs recommend 2–8; examples 4 and 2) decoupled loss optional; rejection sampling token/mask/ratio, upper 5.0; GSPO, M2PO options soft pause; per-token version list
PRIME-RL always async orchestrator.max_off_policy_steps = 8, queue time included; no max_async_level field exists any more ipo (eps 0.1 on absolute probability change) or icepop (0.2 / 5.0) a rollout may span updates; a group shares one dispatch version
NeMo-RL async GRPO max_trajectory_age_steps = 1; in_flight_weight_updates false IS correction required for async; tis, icepop, seq-mask-tis variants pause/resume preserving request state
PipelineRL in-flight after every optimizer step max_lag = null; samples beyond it are counted, not dropped IS-REINFORCE against rollout-server log-probs, clamp 5 never pauses; stale KV kept
slime / Miles streaming, partial rollout, fully-async function slime: none; Miles: --max-weight-staleness (unset = disabled) --use-tis (2.0), --use-rollout-logprobs, --use-opsm (1e-4), custom MIS slime re-queues aborted groups; Miles --pause-generation-mode retract / in_place / abort
OpenRLHF --train.async_enable, queue size 1, partial rollout none; no version tracking found is_correction_* level, mode, gating; threshold [0.5, 5.0] vLLM pause/resume; in-flight samples mix versions
ROLL async_generation_ratio (0 = sync) the same knob; staleness bounded at ceil(ratio) steps train_infer_correction.is_weight (upper 1.2), ratio filters 0.8–1.2; TOPR, TIS, CISPO variants pauses sampling before the update, collects unfinished

Defaults for frameworks other than SkyRL were read on their main branches on 2026-09-17 and are recorded, with commits and the three defaults re-read by hand, in the configuration audit. Within the SkyRL checkout alone there are three documentation drifts: the KV-cache default, the pause semantics, and a generator.use_cache_salt option named in the live docs that does not exist in the June code. Two cross-framework facts matter for benchmark design: only SkyRL, AReaL and Miles bound staleness at admission, the rest drop or wait at the trainer or do not bound it in core; and version provenance is per token in AReaL and verl V1, per sample in PipelineRL, PRIME-RL, NeMo-RL and slime, and absent in OpenRLHF and ROLL.

8. What this means for RL-Infra-Bench

The blueprint already lists "asynchrony and staleness" as a two-GPU, one-node mechanism whose trusted signal is the version-lag distribution plus importance-ratio and clipping statistics, and argues that the magnitude is governed by generation time over update time, in-flight rollouts over batch size, and sync latency over step time. The literature supports that framing and sharpens the observable signature: Applied Compute's formula turns the ratio argument into a prediction a verifier can check, and ESTR's intra/inter split and the missing-old-logits factorization say which quantities must be logged per token. BLUEPRINT.md:716, 796–812

What exists today:

Task shapes the sources make concrete, each with a verifier signal the candidate cannot fake:

  1. Silently off-policy run. Correction is configured but rollout log-probabilities are not attached, so SkyRL's correction path computes nothing. Signal: is_ratio_* metrics absent while max_staleness_steps > 0; the verifier recomputes ratios from logged engine log-probs. Reference: the GLM-5 and Mercor practice of using engine log-probs directly.
  2. Aggregate bound satisfied, per-group bound violated. Long-tail trajectories exceed S while the capacity rule holds, which SkyRL accepts with a warning. Signal: per-group trained_step − scheduled_step distribution against the Applied Compute prediction for the run's C, B, q and ρ. Reference fixes: GLM-5's τ rule, INTELLECT-3's max_off_policy_steps.
  3. Version provenance lost across a pause or resume. Trajectories carry the wrong version after a keep-mode pause or a checkpoint restore. Signal: engine-reported version against trainer version at each step, the same probe the blueprint lists for weight synchronization; the missing-old-logits paper gives the cost envelope for exact recovery.
  4. Stale KV cache across versions. Prefix blocks reused across a sync. Signal: log-prob difference on the shared prefix before and after sync, with the PipelineRL "slightly higher divergence" result as the expected magnitude and Laguna's reset-on-broadcast as the conservative alternative.
  5. Optimizer state across weight pushes. Reproduce GLM-5's optimizer reset as a knob; verify the effect on the importance-ratio distribution, not on final reward.
  6. Mismatch versus staleness attribution. Use the GDN bitwise-parity control as the positive control so that a nonzero ratio is attributable to lag alone, then inject lag through the delay knob.
  7. Staleness ceiling engineering. Given a fixed GPU split, reach a target step time under a staleness bound by choosing C, B and q; the verifier checks measured mean staleness against the closed form and the ESS of the correction weights.

These are candidates, not qualified tasks; each needs the blueprint's positive and negative controls before use. Effects that only appear after thousands of updates, such as the final-reward cost of a given staleness, remain out of scope for the workhorse tier, as the blueprint already states.

9. Evidence boundaries and open questions

Source map

Item Pin
SkyRL staleness manager, capacity formula ../SkyRL/skyrl/train/fully_async_trainer.py:104–176
SkyRL fully-async and correction defaults ../SkyRL/skyrl/train/config/config.py:393–420, 490–520
SkyRL correction math and metrics ../SkyRL/skyrl/backends/skyrl_train/utils/off_policy_correction_utils.py:24–312
SkyRL rollout_is and DPPO losses ../SkyRL/skyrl/backends/skyrl_train/utils/ppo_utils.py:752–863
SkyRL keep-mode pause ../SkyRL/skyrl/backends/skyrl_train/inference_engines/inference_engine_client.py:381–395
SkyRL docs: decomposition, recommendations, TIS vs geometric comparison ../SkyRL/docs/content/docs/algorithms/off_policy_correction.mdx
SkyRL docs: fully async tutorial and config drift ../SkyRL/docs/content/docs/tutorials/fully_async.mdx; configuration/config.mdx:606–618
GLM-5 report §3.2, §3.6, §4.1 arXiv 2602.15763 v2, text extracted from the PDF on 2026-09-17
Qwen formulation paper arXiv 2512.01374 v3
ScaleRL; PipelineRL; AReaL; ROLL Flash; LlamaRL; INTELLECT-3 arXiv 2510.13786; 2509.19128; 2505.24298; 2510.11345; 2505.24034; 2512.16144
DPPO; M2PO; VCPO; ESTR; staleness–LR law; missing old logits arXiv 2602.04879; 2510.01161; 2602.17616; 2607.22186; 2607.01083; 2605.12070
StaleFlow; ECHO-2; Laguna; Kimi K3; DeepSeek-V3.2 arXiv 2601.12784; 2602.02192; 2605.27605; 2607.24653; 2512.02556
Applied Compute; Huang; Mercor; Marin; Hugging Face survey; Wang; Panda blog and issue URLs in the companion notes
Framework configuration audit (verl, AReaL, PRIME-RL, NeMo-RL, PipelineRL, slime, Miles, OpenRLHF, ROLL) analysis/async_rl_staleness/framework_knobs.md, main-branch commits of 2026-09-17
RL-Infra-Bench blueprint and tasks ../RL-Infra-Bench/BLUEPRINT.md:692, 716, 796–812; tasks and tests as linked above
Batch-invariant GDN study ../torchtitan-batch-invarient-GDN/README.md