# Staleness and off-policy drift in asynchronous LLM RL — related-works note Reviewed 2026-09-17. **Staleness is the version lag between the policy that sampled a trajectory and the policy whose gradient that trajectory updates. In asynchronous post-training it is the default condition, not a corner case. The 2025–2026 recipes treat it as a controlled quantity: bounded by admission control or version-drop rules, corrected at the loss with ratios taken against the inference engine's own log-probabilities, and monitored through log-probability-difference and importance-ratio statistics. No source reviewed here calls the problem solved.** Two confounds must be separated before any result is attributed to staleness: training-versus-inference numerical mismatch, which is present at zero lag, and multi-gradient-step mini-batching, which creates staleness with no asynchrony at all. This is a source-pinned reading note, not an execution study. No training runs were launched. Local pins: the SkyRL checkout at `../SkyRL` (commit `152ebec0`, 2026-06-25), the GLM-5 report text extracted from its PDF, and the RL-Infra-Bench blueprint. Per-source records, with which numbers were re-read on the primary page and which come from a delegated sweep, are in four companion notes: [systems](analysis/async_rl_staleness/systems_sources.md), [estimators](analysis/async_rl_staleness/estimator_sources.md), [mismatch](analysis/async_rl_staleness/mismatch_sources.md) and [the 2026 recency sweep](analysis/async_rl_staleness/recency_2026.md). ## 1. Definitions and the decomposition Let μ be the sampling distribution the inference engine actually realized, π_old the trainer's recomputation of the same weights, and π_θ the policy being updated. The token-level importance ratio an on-policy surrogate needs factors as ```text π_θ(x_t) / μ_old(x_t) = [ π_old(x_t) / μ_old(x_t) ] × [ π_θ(x_t) / π_old(x_t) ] training–inference mismatch policy staleness ``` The decomposition appears in the Qwen team's formulation paper, in SkyRL's off-policy-correction documentation (crediting verl's rollout-correction math note), and as r = r_d·r_s in the May 2026 "missing old logits" paper. The first factor is nonzero in a perfectly synchronous run because inference kernels, precision, parallelism and MoE routing differ from the trainer's. The second is what asynchrony, in-flight updates, partial rollouts and mini-batching add. The Qwen paper's first-order argument is that the token-level surrogate tracks the true sequence-level objective only while both factors stay small. [Qwen, arXiv 2512.01374](https://arxiv.org/abs/2512.01374) · [SkyRL docs](../SkyRL/docs/content/docs/algorithms/off_policy_correction.mdx) · [Guan et al., arXiv 2605.12070](https://arxiv.org/abs/2605.12070) **Staleness is counted in policy versions, and the accounting differs by system.** SkyRL counts the training step that consumes a group minus the step at which it was scheduled. AReaL, ROLL Flash and StaleFlow count from the version that started generation. GLM-5 records the ordered list of versions that touched a response and measures from the oldest. ESTR separates intra-trajectory staleness (versions spanned inside one rollout under in-flight updates) from inter-trajectory lag (last token's version to the target policy). Applied Compute splits mean staleness into a pre-queue part accrued during generation and an in-queue part accrued while waiting for the trainer. Numbers from different papers are therefore not comparable without restating the definition. [ESTR](https://arxiv.org/abs/2607.22186) · [Applied Compute](https://www.appliedcompute.com/research/staleness-in-fully-async-rl) **Under in-flight updates the behavior policy is not a checkpoint.** GLM-5 states that tracking exact behavior probabilities across the versions a single response spans would require keeping every intermediate checkpoint, and instead uses the rollout engine's logged log-probabilities as the behavior proxy. The missing-old-logits paper measures what recovering π_old exactly would cost (snapshots: 40–76 GB and 25–150 s per step on 4B and 30B models) and proposes an exponentially weighted parameter average as the proxy. [GLM-5 §4.1.2, arXiv 2602.15763](https://arxiv.org/abs/2602.15763) ## 2. Where staleness comes from | Source | Mechanism | Pinned example | |---|---|---| | Decoupled generation and training | Engines generate continuously; the trainer consumes whatever finished; weights are pushed back periodically | GLM-5 pushes every K gradient updates and resets the optimizer after each push; Laguna broadcasts every 2 optimizer steps; ECHO-2 publishes every S−1 steps | | In-flight weight update | An unfinished sequence continues under new weights | PipelineRL keeps the stale KV cache; AReaL recomputes it; INTELLECT-3 lets one trajectory span several policies | | Partial rollouts | Long trajectories are paused at a budget and resumed next iteration | Kimi k1.5/K2/K3 pause once a fraction λ completes; APRIL resumes over-provisioned rollouts across up to five versions with no correction | | Long-tail trajectories | A group scheduled at step i finishes after several steps | SkyRL's capacity rule bounds aggregate lag, not per-group lag; violators are accepted with a warning | | Mini-batching and batch reuse | Several gradient steps per sampled batch make later steps off-policy with no asynchrony | Qwen varies N ∈ {1,2,4,8} mini-batches; μ-GRPO and PNPO reuse batches for hundreds of updates | | Stale KV cache | Prefix blocks computed under an older policy are reused after a sync | SkyRL `clear_kv_cache_on_weight_sync` defaults to false in code; Laguna resets the cache on every broadcast | | Text round-trips | Re-tokenizing engine text breaks token-level correspondence and masquerades as off-policyness | GLM-5 token-in-token-out gateway; Mercor rewrote its harness to use `/completions` | ## 3. How systems bound staleness **Admission control.** AReaL rejects new generation requests when ⌊(N_r−1)/B⌋ > i+η, with N_r trajectories generated so far, B the batch size and i the current version; η = 8 for math and 4 for code. SkyRL reimplements this as a capacity rule, `accepted + running < (max_staleness_steps + current_step) × mini_batch_size`, with a buffer of `mini_batch_size × (S+1)` finished groups and the constraint `mini_batch_size ≤ workers ≤ mini_batch_size × (S+1)`; default S = 4. ROLL Flash caps its sample buffer at (1+α)×batch so no generated sample can violate the bound. StaleFlow reserves a version tag at admission and consumes a trajectory only if V_traj + η ≥ V_buf. All four are aggregate or per-trajectory bounds on versions, not on wall-clock. [AReaL](https://arxiv.org/abs/2505.24298) · [fully_async_trainer.py:104–176](../SkyRL/skyrl/train/fully_async_trainer.py) · [ROLL Flash](https://arxiv.org/abs/2510.11345) · [StaleFlow](https://arxiv.org/abs/2601.12784) **Drop rules at training time.** GLM-5 discards a response when the current version minus its oldest version exceeds τ (unreported). INTELLECT-3 discards rollouts generated by more than `max_off_policy_steps = 8` policies. ECHO-2 discards below v_t − S. verl's V1 sampler drops, or optionally waits on, prompt groups at `max_off_policy_threshold = 8` versions; NeMo-RL evicts trajectories older than `max_trajectory_age_steps` (default 1); PipelineRL's source counts samples beyond `max_lag` rather than dropping them, although the March 2026 Hugging Face survey classifies it as version rejection. SkyRL's documentation prefers proactive admission to over-generating and dropping, on the grounds of wasted compute. **In-flight handling is where implementations diverge most.** PipelineRL continues in-progress sequences with the stale KV cache and reports only slightly higher divergence than recomputing it. AReaL discards and recomputes the cache. The June SkyRL checkout freezes requests with vLLM's keep-mode pause and resumes them after the sync; its own tutorial and the Hugging Face survey still describe an abort-and-refeed loop, so the behavior changed between March and June 2026. Laguna blocks any in-flight rollout step on an update so no step straddles a refresh, and resets the KV cache. ECHO-2 workers simply switch version when a snapshot finishes installing. The Prime Intellect 1T run drops by `max_off_policy_steps` and salts the KV cache by version. [PipelineRL](https://arxiv.org/abs/2509.19128) · [inference_engine_client.py:381–395](../SkyRL/skyrl/backends/skyrl_train/inference_engines/inference_engine_client.py) · [Laguna, arXiv 2605.27605](https://arxiv.org/abs/2605.27605) **What sets the mean staleness.** Applied Compute derives closed forms in five parameters: trainer utilization ρ, batch size B in rollouts, sampling concurrency C, queue factor q and a response-length tailness multiplier. Rollout-bound: staleness ≈ C·M_tail/B + ρ. Train-bound: C·M_tail/(ρB) + (2q+ρ−1)/(2ρ). Predictions on real Qwen3-8B GRPO runs were within 0.06–0.27 versions of measurement; setting q = 1 cut staleness 35% at unchanged step time in one configuration. This is the same dimensionless-ratio framing the RL-Infra-Bench blueprint uses, now with a validated formula. [Applied Compute, 2026-07-03](https://www.appliedcompute.com/research/staleness-in-fully-async-rl) **Production ceilings.** Mercor's SkyRL runs use S = 3 with mini-batch 16 and 16 samples per prompt, a 1,024-trajectory ceiling, run at 550 (35B) and 300 (397B) concurrent rollouts. Laguna caps at 10 optimizer steps and never reaches it. GLM-5 runs more than 1k concurrent rollouts with τ and K unreported. Marin's 67B runs recommend a staleness allowance of 4 updates. [Mercor guide, 2026-09-01](https://www.mercor.com/blog/training-frontier-knowledge-work-agents-a-397b-rl-training-guide-with-skyrl/) ## 4. Loss-level corrections All of these take the ratio against the rollout engine's logged log-probabilities. They differ in what they do with it. | Correction | Rule | Where pinned | |---|---|---| | Truncated importance sampling, token | multiply the loss by `min(π/μ, C)` per token; SkyRL default C = 2.0; verl default threshold 2.0; PipelineRL clamps at 5; LlamaRL's AIPO uses C in [2, 10] | SkyRL `tis_ratio_type="token"`; Yao et al. | | Truncated importance sampling, sequence | same with the product over tokens; SkyRL default C = 5.0; Qwen uses cap 5; VCPO uses 8 | SkyRL `tis_ratio_type="sequence"`; Liu et al. | | Sequence mask, geometric or product | zero a sequence when the geometric mean of ratios leaves [0.99, 1.01] or the product leaves [0.5, 2.0] | SkyRL `sequence_mask_metric`; verl MIS mode; DeepSeek-V3.2 masks negative-advantage sequences whose mean log-ratio exceeds δ | | Outlier-token sequence mask | zero a sequence if any token ratio is outside [1e-4, 100]; verl vetoes below 1e-4 | SkyRL `outlier_token_is_threshold_*` | | Token mask (IcePop family) | zero individual tokens with ratio outside [1/β, β]; Ring-1T and INTELLECT-3 use [0.5, 5]; GLM-5 synchronous RL uses β = 2 | GLM-5 §3.2; SkyRL `token_mask_is_threshold_*` | | Direct double-sided IS (GLM-5 agentic loss) | r_t = π_θ/π_rollout with no π_old; gradient `f(r_t) Â log π_θ`, f zero outside [1−ε_l, 1+ε_h] | GLM-5 §4.1.2 eq. 3–5; SkyRL `policy_loss_type="rollout_is"` | | CISPO | clip the weight, not the update, and keep the gradient through log π; Laguna's asymmetric clip gives an effective range [0, 5] | MiniMax-M1; ScaleRL; Laguna | | DPPO | binary mask on the probability difference π_θ − μ (total variation) or a binary KL, thresholds 0.2 or 0.05, rollout log-probs as behavior policy | [arXiv 2602.04879](https://arxiv.org/abs/2602.04879); SkyRL `policy_loss_type="dppo"`; Mercor | | Second-moment trust region (M2PO) | no ε-clip; mask the highest (log r)² tokens until the batch mean of (log r)² is below 0.04 | [arXiv 2510.01161](https://arxiv.org/abs/2510.01161) | | Entropy-scaled trust region (ESTR) | keep a token iff δ_t²/(H_t+ε) ≤ τ with H_t the behavior entropy | [arXiv 2607.22186](https://arxiv.org/abs/2607.22186) | | Three-policy objectives | PPO-clip π_θ/π_prox times π_prox/π_behav (AReaL decoupled PPO); verl decoupled mode; A-3PO interpolates π_prox instead of recomputing it | AReaL; verl docs | | Optimizer-side | scale the learning rate by √(ESS_off/ESS_on) (VCPO); project or skip updates by gradient cosine (GAC); S·η_max ≈ 1.6e-6 as an empirical law | [VCPO](https://arxiv.org/abs/2602.17616); [arXiv 2607.01083](https://arxiv.org/abs/2607.01083) | SkyRL applies its masks multiplicatively in a fixed order (outlier sequence mask, token mask, sequence mask) and logs `is_ratio_{mean,std,max,min}`, clip fractions and `rollout_train_logprobs_abs_diff_*`. Every correction is silently disabled when no rollout log-probabilities are attached to the batch. [off_policy_correction_utils.py:257–312](../SkyRL/skyrl/backends/skyrl_train/utils/off_policy_correction_utils.py) ## 5. How much staleness is tolerable: reported evidence | Study | Setup | Finding | Boundary | |---|---|---|---| | Asynchronous RLHF (Oct 2024) | Pythia 410M–2.8B; N ∈ {1..64} mini-batch updates per generation | online DPO still learns at N=64; PPO and RLOO degrade sharply; larger policies more robust | production setting is one-step lag only | | AReaL (May 2025) | 1.5B, AIME24, η ∈ {0,1,2,4,8,∞} | naive PPO 42.0 → 23.3 at η=4; decoupled PPO 42.0 / 42.1 / 41.8 / 42.2 / 41.0 / 36.9; about 2.1× throughput at η=4 | one model, one benchmark | | ScaleRL (Oct 2025) | 8B dense and 17B×16 MoE; PipelineRL-k, k ∈ {1,8} | k=8 chosen; same asymptote as PPO-off-policy-8 with better compute efficiency; FP32 head raises asymptote 0.52 → 0.61 | two k values; the precision result is a mismatch effect | | ROLL Flash (Oct 2025) | Qwen3-8B-Base; α sweep | α=2 optimal for most settings; decoupled PPO, TIS, CISPO, TOPR and vanilla GRPO comparable at α=2 and α=8 | curves only | | M2PO (Oct 2025) | 1.7B–32B; data delayed by k ∈ {0..256} updates | Qwen2.5-Math-7B: GRPO 49.3 on-policy, 45.7 at k=256; GSPO 44.7; M2PO 48.8; decoupled PPO, TOPR, CISPO break down early at k=256 | staleness simulated, not a live pipeline | | Qwen "Stabilizing RL" (Dec 2025) | Qwen3-30B-A3B, FP8 inference, N ∈ {1,2,4,8} mini-batches | N=2 needs routing replay plus clipping; N ≥ 4 needs full R3 plus clipping or collapses | staleness from mini-batching; math only | | ECHO-2 (Feb 2026) | Qwen3-8B; S ∈ {3,4,6,11} | S ≤ 6 within about 5% of synchronous; S = 11 diverges | consumer-GPU, wide-area regime | | VCPO (Feb 2026) | Qwen2.5-7B; PipelineRL-k with k = 10–12, stress to 128 | MATH-500 72.0 sync vs 71.6 at k=10 with 1.45× fewer GPU-hours; TIS, MIS, M2PO collapse at k=12 on GSM8K | ≤7B | | StaleFlow (Jan/Aug 2026) | Qwen family to 30B-A3B; η ∈ {0,1,3}, 10 | η=3 matches synchronous convergence; η=10 or no control collapses | single η in the convergence check | | Huang (May 2026) | toy bit-sequence MDP, horizon 1024 | every method collapses at K=12; sequence-level TIS matches sync from batch 32 and exceeds it at 128 while token-level keeps collapsing | synthetic | | Staleness–LR law (Jul 2026) | Llama-3.2-1B/3B, S ∈ {8,16,32} | η_max ∝ 1/S with S·η_max ≈ 1.6e-6; collapse time t·η ≈ 3.2e-5 regardless of S | tiny models; local analysis | | ESTR (Jul 2026) | Qwen3-30B, Qwen2.5-7B; Δ_intra to 9, Δ_inter to 30 | IcePop and KPop collapse within a few hundred steps on multi-turn GSM8K; ESTR 95.69; 2.6× throughput | τ not captured; no TIS/MIS baselines | | Marin standup (Sep 2026) | 67B; staleness 0/1/2/4 updates | 1: −30% wall-clock, +0.3 quality (5 seeds); 4: −55%, −2.6 (3 seeds), judged non-inferior | informal; acceptable range "not established" per the companion issue | | Mercor/SkyRL (Sep 2026) | 35B and 397B; S=3 | DPPO within noise of the GLM-5 loss at epoch 1 but shorter, more numerous turns (21→32 turns; 834→588 tokens per turn); mean log-prob diff under 0.03 treated as healthy | one run per size | The pattern: tolerated staleness moved from about one step in 2024 to 3–8 versions in 2025–2026 production recipes, and to 50–256 updates in method papers on small models, but only after ratios were taken against engine log-probabilities, masks were added, and MoE routing was replayed. Quote a tolerance only together with its loss, its mismatch controls and its staleness definition. ## 6. The mismatch confound Mismatch is measurable at zero staleness and is often the larger factor. Measured magnitudes: KL(engine‖trainer) roughly 1e-4–1e-3 in bf16 on dense models, 1e-2 in FP8, up to 1e-1 for MoE; a 30B-A3B MoE flips at least one expert on 94% of tokens between SGLang and Megatron. Fixes and costs: - **Precision**: FP32 LM-head logits (ScaleRL 0.52 → 0.61; MiniMax-M1 correlation 0.9x → 0.99x). - **Routing replay** (R3, Keep Routing): train/infer KL 1.535e-3 → 7.5e-4 on Qwen3-30B-A3B at under 3% rollout latency; adds bias under mini-batching because routing is frozen across mini-batches. - **Deterministic operators**: GLM-5 replaced a non-deterministic top-k in its sparse-attention indexer with `torch.topk` after the alternatives collapsed within a few steps. - **Batch-invariant or bitwise-consistent kernels**: 20–60% sampling-throughput cost and 2–5× trainer cost depending on architecture; the vLLM+TorchTitan bitwise stack is 2.4× slower; IsoExec reports 25% end-to-end overhead. Two August 2026 studies test whether removing mismatch helps under staleness. Wang's GDN study reaches an exactly zero gap at lag 0 with batch-invariant kernels, then sweeps off-policy windows of 4, 12 and 32: at window 12 on MATH the parity arm peaks near 0.7 train reward against 0.6 for vLLM native, but at window 4 the arms interleave, on Search-R1 they are indistinguishable, and at window 32 both degenerate. The author's conclusion is to use parity as a debugging control, not a production default. Panda's MoE study gets K3 to zero on Qwen3.6-35B-A3B and raises held-out solve from 63.9% to 77.4% synchronously, but an asynchronous run with stale weights and cache and no mitigation falls to 59.4%, recovering to 76.6% with router recall plus CISPO. Both confirm the ordering in §1: fix numerics first so that any residual ratio is attributable to lag, then bound and correct the lag. [Wang, Aug 2026](https://yichuan-w.github.io/blog/GDN-train-inference-mismatch-asyncRL/) · [Panda, 2026-08-17](https://kiddyboots216.github.io/mismatch/) · [`../torchtitan-batch-invarient-GDN`](../torchtitan-batch-invarient-GDN/README.md) ## 7. Framework knobs | Framework | Async mode | Staleness knob and default | Correction options | In-flight handling | |---|---|---|---|---| | SkyRL (local, June 2026) | `trainer.fully_async` | `max_staleness_steps` = 4; `num_parallel_generation_workers` = 768; `clear_kv_cache_on_weight_sync` false in code, "true" in the config page | none by default; `off_policy_correction.*`, `rollout_is`, `dppo`, `cispo`, `gspo`, `sapo` | keep-mode pause; requests frozen and resumed | | verl | V1 trainer (`sync`, `colocate_async`, `separate_async`); fully-async recipe; one-step-off | V1 `trainer.v1.sampler.max_off_policy_threshold` = 8 with `drop` or `wait`; fully-async `async_training.staleness_threshold` = 0 (a fraction; keep below 1) | `algorithm.rollout_correction`: TIS or rejection × token, sequence, geometric; threshold 2.0; `bypass_mode`; IcePop preset | V1 keeps tokens and log-probs, rebuilds KV, resumes on the new version | | AReaL | native async | `rollout.max_head_offpolicyness` = 0 (docs recommend 2–8; examples 4 and 2) | decoupled loss optional; rejection sampling token/mask/ratio, upper 5.0; GSPO, M2PO options | soft pause; per-token version list | | PRIME-RL | always async | `orchestrator.max_off_policy_steps` = 8, queue time included; no `max_async_level` field exists any more | `ipo` (eps 0.1 on absolute probability change) or `icepop` (0.2 / 5.0) | a rollout may span updates; a group shares one dispatch version | | NeMo-RL | async GRPO | `max_trajectory_age_steps` = 1; `in_flight_weight_updates` false | IS correction required for async; tis, icepop, seq-mask-tis variants | pause/resume preserving request state | | PipelineRL | in-flight after every optimizer step | `max_lag` = null; samples beyond it are counted, not dropped | IS-REINFORCE against rollout-server log-probs, clamp 5 | never pauses; stale KV kept | | slime / Miles | streaming, partial rollout, fully-async function | slime: none; Miles: `--max-weight-staleness` (unset = disabled) | `--use-tis` (2.0), `--use-rollout-logprobs`, `--use-opsm` (1e-4), custom MIS | slime re-queues aborted groups; Miles `--pause-generation-mode` retract / in_place / abort | | OpenRLHF | `--train.async_enable`, queue size 1, partial rollout | none; no version tracking found | `is_correction_*` level, mode, gating; threshold [0.5, 5.0] | vLLM pause/resume; in-flight samples mix versions | | ROLL | `async_generation_ratio` (0 = sync) | the same knob; staleness bounded at ceil(ratio) steps | `train_infer_correction.is_weight` (upper 1.2), ratio filters 0.8–1.2; TOPR, TIS, CISPO variants | pauses sampling before the update, collects unfinished | Defaults for frameworks other than SkyRL were read on their `main` branches on 2026-09-17 and are recorded, with commits and the three defaults re-read by hand, in the [configuration audit](analysis/async_rl_staleness/framework_knobs.md). Within the SkyRL checkout alone there are three documentation drifts: the KV-cache default, the pause semantics, and a `generator.use_cache_salt` option named in the live docs that does not exist in the June code. Two cross-framework facts matter for benchmark design: only SkyRL, AReaL and Miles bound staleness at admission, the rest drop or wait at the trainer or do not bound it in core; and version provenance is per token in AReaL and verl V1, per sample in PipelineRL, PRIME-RL, NeMo-RL and slime, and absent in OpenRLHF and ROLL. ## 8. What this means for RL-Infra-Bench The blueprint already lists "asynchrony and staleness" as a two-GPU, one-node mechanism whose trusted signal is the version-lag distribution plus importance-ratio and clipping statistics, and argues that the magnitude is governed by generation time over update time, in-flight rollouts over batch size, and sync latency over step time. The literature supports that framing and sharpens the observable signature: Applied Compute's formula turns the ratio argument into a prediction a verifier can check, and ESTR's intra/inter split and the missing-old-logits factorization say which quantities must be logged per token. [BLUEPRINT.md:716, 796–812](../RL-Infra-Bench/BLUEPRINT.md) What exists today: - `train-five-steps` enables token TIS with cap 2.0, requests rollout log-probabilities, and verifies policy versions 1–5 plus evaluation version 6. [solve.sh:184](../RL-Infra-Bench/tasks/skyrl-vllm/train-five-steps/solution/solve.sh) · [verify.py:48](../RL-Infra-Bench/tasks/skyrl-vllm/train-five-steps/tests/verify.py) - The submitted-job qualification rejects reports whose last trial carries a stale policy version. [test_submitted_rl.py:589](../RL-Infra-Bench/evals/modal-workspace-jobs/tests/test_submitted_rl.py) - `gdn-parity` covers the mismatch confound with exact log-probability comparison. - SkyRL's fully asynchronous Harbor entrypoint is a listed source candidate with no task yet. [README.md:145](../RL-Infra-Bench/README.md) Task shapes the sources make concrete, each with a verifier signal the candidate cannot fake: 1. **Silently off-policy run.** Correction is configured but rollout log-probabilities are not attached, so SkyRL's correction path computes nothing. Signal: `is_ratio_*` metrics absent while `max_staleness_steps > 0`; the verifier recomputes ratios from logged engine log-probs. Reference: the GLM-5 and Mercor practice of using engine log-probs directly. 2. **Aggregate bound satisfied, per-group bound violated.** Long-tail trajectories exceed S while the capacity rule holds, which SkyRL accepts with a warning. Signal: per-group `trained_step − scheduled_step` distribution against the Applied Compute prediction for the run's C, B, q and ρ. Reference fixes: GLM-5's τ rule, INTELLECT-3's `max_off_policy_steps`. 3. **Version provenance lost across a pause or resume.** Trajectories carry the wrong version after a keep-mode pause or a checkpoint restore. Signal: engine-reported version against trainer version at each step, the same probe the blueprint lists for weight synchronization; the missing-old-logits paper gives the cost envelope for exact recovery. 4. **Stale KV cache across versions.** Prefix blocks reused across a sync. Signal: log-prob difference on the shared prefix before and after sync, with the PipelineRL "slightly higher divergence" result as the expected magnitude and Laguna's reset-on-broadcast as the conservative alternative. 5. **Optimizer state across weight pushes.** Reproduce GLM-5's optimizer reset as a knob; verify the effect on the importance-ratio distribution, not on final reward. 6. **Mismatch versus staleness attribution.** Use the GDN bitwise-parity control as the positive control so that a nonzero ratio is attributable to lag alone, then inject lag through the delay knob. 7. **Staleness ceiling engineering.** Given a fixed GPU split, reach a target step time under a staleness bound by choosing C, B and q; the verifier checks measured mean staleness against the closed form and the ESS of the correction weights. These are candidates, not qualified tasks; each needs the blueprint's positive and negative controls before use. Effects that only appear after thousands of updates, such as the final-reward cost of a given staleness, remain out of scope for the workhorse tier, as the blueprint already states. ## 9. Evidence boundaries and open questions - Production accounts (GLM-5, Laguna, Mercor, DeepSeek, Kimi K3) report settings without a staleness sweep; the sweeps that exist use one model family each, and several induce staleness by delayed data or mini-batching rather than a live pipeline. - "Tolerable staleness" is recipe-dependent and definition-dependent; the ordering of methods flips between studies (M2PO beats GSPO at k=256; ROLL Flash finds them comparable at α ≤ 8; VCPO finds M2PO collapsing at k=12 on GSM8K). - Applied Compute predicts staleness but not its effect on reward; Huang's collapse threshold is on a toy MDP; the scaling law is on 1B–3B models. Nobody has published a reward-versus-staleness curve at frontier scale. - Whether reusing a stale KV cache is harmless at scale rests on PipelineRL's divergence measurement at 7B and on defaults elsewhere; Laguna chose the opposite. - Documentation drifts inside a single SkyRL checkout show that framework defaults must be pinned to a revision, which is the same rule this workspace applies to benchmarks. ## Source map | Item | Pin | |---|---| | SkyRL staleness manager, capacity formula | `../SkyRL/skyrl/train/fully_async_trainer.py:104–176` | | SkyRL fully-async and correction defaults | `../SkyRL/skyrl/train/config/config.py:393–420, 490–520` | | SkyRL correction math and metrics | `../SkyRL/skyrl/backends/skyrl_train/utils/off_policy_correction_utils.py:24–312` | | SkyRL rollout_is and DPPO losses | `../SkyRL/skyrl/backends/skyrl_train/utils/ppo_utils.py:752–863` | | SkyRL keep-mode pause | `../SkyRL/skyrl/backends/skyrl_train/inference_engines/inference_engine_client.py:381–395` | | SkyRL docs: decomposition, recommendations, TIS vs geometric comparison | `../SkyRL/docs/content/docs/algorithms/off_policy_correction.mdx` | | SkyRL docs: fully async tutorial and config drift | `../SkyRL/docs/content/docs/tutorials/fully_async.mdx`; `configuration/config.mdx:606–618` | | GLM-5 report §3.2, §3.6, §4.1 | arXiv 2602.15763 v2, text extracted from the PDF on 2026-09-17 | | Qwen formulation paper | arXiv 2512.01374 v3 | | ScaleRL; PipelineRL; AReaL; ROLL Flash; LlamaRL; INTELLECT-3 | arXiv 2510.13786; 2509.19128; 2505.24298; 2510.11345; 2505.24034; 2512.16144 | | DPPO; M2PO; VCPO; ESTR; staleness–LR law; missing old logits | arXiv 2602.04879; 2510.01161; 2602.17616; 2607.22186; 2607.01083; 2605.12070 | | StaleFlow; ECHO-2; Laguna; Kimi K3; DeepSeek-V3.2 | arXiv 2601.12784; 2602.02192; 2605.27605; 2607.24653; 2512.02556 | | Applied Compute; Huang; Mercor; Marin; Hugging Face survey; Wang; Panda | blog and issue URLs in the companion notes | | Framework configuration audit (verl, AReaL, PRIME-RL, NeMo-RL, PipelineRL, slime, Miles, OpenRLHF, ROLL) | `analysis/async_rl_staleness/framework_knobs.md`, main-branch commits of 2026-09-17 | | RL-Infra-Bench blueprint and tasks | `../RL-Infra-Bench/BLUEPRINT.md:692, 716, 796–812`; tasks and tests as linked above | | Batch-invariant GDN study | `../torchtitan-batch-invarient-GDN/README.md` |