Production LLM pipelines: what the sources say, 21 September 2026

This note backs the "life cycles" figure, the code-by-stage band and the RL-loop anatomy. It merges the companion notes written earlier the same day (../../llm_infra_review/production_lifecycle_sources_v2.md, rl_pipeline_sources_v2.md; tag note-2026-09-21), three delegated sweeps (tag agent-reported), and eight primary pages re-read in this session (tag spot-checked): the slime README, GLM-4.5 §3.5, the Olmo 3 blog, the checkpoint-engine README, the Terminal-Bench 2.1 news post and task packages, and the PostTrainBench and InferenceBench READMEs. Earlier workspace notes are prior-note. Numbers are in data/pipelines.json and data/rl_loop.json; re-read agent-reported entries before quoting them in a paper.

1. The stage vocabulary is stable even where the names differ

Source Data Pretrain Midtrain Post-train Agentic RL Serve
Llama 3 (report) curation, dedup, mixture 16K H100, 4D parallelism long-context continuation, annealing rejection sampling, SFT, DPO rounds — inference optimization, FP8
DeepSeek-V3 / R1 (report; serving code) 14.8T tokens 2,048 H800, DualPipe, FP8 32K → 128K (YaRN) cold start → reasoning RL → rejection sampling + SFT → RL for all scenarios — prefill unit 4 nodes, decode unit 40 nodes; DeepEP/DeepGEMM/EPLB open
Qwen3 (report) > 30T tokens General S1, Reasoning S2 Long Context stage long-CoT cold start → reasoning RL → thinking-mode fusion → general RL Qwen3.5: disaggregated async RL (blog only) —
Kimi K2 / K3 (report; two components open) 15.5T tokens MuonClip, zero loss spike 4K → 128K; K3 to 1M SFT, joint RL; checkpoint-engine (1T update < 30 s) agentic data synthesis, partial rollouts; AgentENV microVMs K3 §5.4
GLM-4.5 (report; RL on open slime) corpus 355B-A32B repo-level code · synthetic reasoning · long-context & agent expert training → unified training; SFT, reasoning/agentic/general RL on slime disaggregated async pools, trajectory pool, Docker runtime SGLang commands
GLM-5 (report; slime) 27T tokens 744B-A40B 4K → 128K → 200K SFT → reasoning RL → general RL → cross-stage distillation fully decoupled, > 1k concurrent rollouts, heartbeat fault tolerance PD disaggregation in the loop
Llama 4 (blog) — > 30T tokens, FP8, 32K GPUs mid-training for 10M context lightweight SFT → online RL → lightweight DPO fully async online RL framework —
DeepSeek-V4 (report) — 32–33T tokens, FP4 experts 1M context domain specialists → on-policy distillation preemptible generation service (token WAL); DSec sandboxes on-disk KV cache
MiniMax-M2 (report) — 29.2T tokens 8K → 32K → 192K SFT, CISPO RL Forge: gateway, data pool, windowed FIFO —
Nemotron 3 Super (report + open recipe) data-prep (subset) Megatron-Bridge, 25T tokens LC-phase CPT at 1M on GB200 SFT → RLVR/SWE-RL/RLHF → MTP healing (NeMo-RL) SWE-RL, NeMo-Gym, 4,000 env instances per batch quantization, evaluator
OLMo 3 (fully open) Dolma 3 tools OLMo-core, 5.9T, up to 1,024 H100 Dolmino 100B; Longmino ~50B Open Instruct: SFT → DPO → RLVR (OlmoRL async) — —
SmolLM3 (fully open) datatrove nanotron, 3 stages to 11.1T, 384 H100 long-context 100B; reasoning mid-training 140B TRL: SFT → APO, merging — —
Apertus (open pretraining) pretrain-data Megatron fork, up to 4,096 GPUs 4K → 64K SFT → QRPO — —
LLM360 K2 / K2-V2 (fully open) TxT360 480 A100, weekly hardware failures four mid-training stages to 512K SFT — —
Marin (open lab) data browser, executor DAG Levanter on TPU, 12.7T (8B) — SFT; RL by 2026 — —
Arcee Trinity Large (report; RL on prime-rl) — 2,048 B300, modified TorchTitan ~117B long-context SFT at 64K; RL with prime-rl — —
Mercor 397B (open recipe) — — — — SkyRL async, S=3, 300–550 concurrent rollouts vLLM rollout engines

Observations that shape the page:

2. The RL loop, and where the implementations differ

The sweep decomposes every stack into six roles: a trainer (Megatron-Core in slime, AReaL, SkyRL/Mercor, NeMo-RL, ROLL and verl; FSDP2 in prime-rl, torchforge and OpenRLHF; hidden behind LoRA in Tinker), a rollout service with an OpenAI-style endpoint (vLLM or SGLang), a weight-sync path (NCCL broadcast in verl, SkyRL, Magistral, TRL and NeMo-RL; a checkpoint engine in Kimi at about 20 s per 1T model; HTTP relays in prime-rl; delta patches in slime; an RDMA store in torchforge), a queue or buffer (slime Data Buffer, verl MessageQueue and TransferQueue, NeMo-RL replay buffer, LongCat SampleQueue, MiniMax Data Pool), reward/verifier and sandbox services scaled on CPU (Prime Sandboxes > 4,000, LongCat 32,000 environments, Kimi K2.5 up to 100,000 tasks, DeepSeek DSec "hundreds of thousands"), and an orchestrator (Ray almost everywhere; Monarch in torchforge).

They differ in placement (colocated hybrid engines for reasoning RL in Kimi K2, Seed1.5-Thinking, verl and GLM-4.5's synchronous mode; disaggregated GPUs for agentic RL in GLM-5, INTELLECT-3's 16:44 split, Mercor's 12:8, NeMo-RL's async mode, RollArt), in synchrony (bounded-staleness knobs: AReaL η = 4/8, INTELLECT-3 max_off_policy_steps = 8, Mercor 4, NeMo-RL max_trajectory_age_steps = 1, verl staleness_threshold, SkyRL max_staleness_steps; in-flight updates and partial rollouts; 2–4× speedups reported), and in the train–inference mismatch fix (importance-sampling correction by default in TRL and required in NeMo-RL; FP32 LM head in MiniMax; batch-invariant kernels in Thinking Machines and TitanRL). This matches the workspace's staleness note (../../ASYNC_RL_STALENESS.md).

The five-part conceptual loop on the page (prompts & environments → rollout → score & admit → learn → publish weights, with checkpoint/recovery and placement/networking/sandbox rails) is the earlier companion note's proposal, labeled as a synthesis illustrated with slime and GLM-4.5. It is not any repository's class diagram.

3. What the audited tasks touch, per component

4. Points to settle with the user

  1. Canonical case study. The page uses GLM-4.5 + slime as the worked example because the report and the code line up (report §3.5 describes exactly the three modules the README names). Nemotron 3 Super is the most complete open recipe; OLMo 3 is the most complete open pipeline. Any of the three could lead.
  2. Whether to treat "Serve" as one stage or two. Rollout engines inside RL and user-facing serving share code but not operational contracts. The page keeps one column and notes the dual role; splitting it would give agentic RL a cleaner infrastructure story.
  3. Family entries. PostTrainBench (28), InferenceBench (4), WeirdML (11) and EdgeBench (51 public) count once each. Counting instances would triple the post-training column without adding verifier diversity.
  4. Agent-reported numbers (DeepSeek-V4 sandbox counts, Kimi K2.5 100,000 concurrent tasks, LongCat 32,000 environments, Qwen3.5 claims) are on the page with dashed tags. They should be re-read before any of them reaches a paper.
  5. Supplemental benchmarks (ISO-Bench, FlashInfer-Bench, CommBench, ATE-Bench, KernelBench) are rings, not counts. Promoting any of them to a primary row needs a task-level read.

5. Revision 2.1 (22 September 2026)

Applied from the user's review of the page: prose now runs at full width; six older or thinner reports (Llama 3, DeepSeek-V3/R1, Llama 4, Apertus, LLM360 K2, Arcee Trinity) are hidden behind a toggle and dimmed when shown; Mercor stays visible. Continued pretraining, annealing and long-context extension were merged into the pretraining column because they share the objective and trainer and no audited task targets them separately. On-policy distillation joined the post-training vocabulary with four code chips (NeMo-RL's on-policy and X-token distillation recipes, TRL's GKDTrainer, Thinking Machines' Tinker post, Xiaomi's MOPD/MOPD2 paper) after re-reading each source. The Marin row now cites the Open Athena 535B launch note (August 2026 start on CoreWeave GPUs with JAX/XLA:GPU, 18T tokens, MFU 21% → 24%, preregistered loss predictions; Snowball 67B-A2B in post-training). The RL loop moved to two deep-dive pages: posttrain.html (recipes as sequences for all 18 pipelines, frameworks by sub-stage, twelve stabilizers with sources) and agentic.html (the loop with 59 implementations, ten fleets, twelve staleness entries, ten ledger rows, seven RL-Infra-Bench families). The grid no longer highlights the agentic column.

6. Revision 2.2 (22 September 2026)

The user asked that the pages read as a review rather than as the case for a benchmark under construction. Every mention of RL-Infra-Bench on the pages was removed: the coverage figure's target box and key item, the loop drawer's "families" line, and the agentic page's blueprint-derived table, which is now "Mechanisms that only exist across devices" with three columns (mechanism, why it needs several devices, what documents it). The overview's post-training section became three blocks: the RL loop (new page loop.html, with the figure, drawer and a component table), post-training recipes, and agentic RL infrastructure. The mapping from incidents to benchmark mechanism families remains in this note's §3 and in mimo_v26_review.md §8 for the benchmark work, off the pages.

7. Revision 2.3 (22 September 2026)

After the user's review of the published draft: the coverage figure moved to the top of the overview, the title became "AI-Infra Environment Coverage" with the question as a subtitle, every mono subhead was removed, captions and section text were cut to about half (overview prose from roughly 700 to 480 words), and the changelog became a working-copy appendix that deploy_site.py strips before publishing. Scale's RSI Bench taxonomy (ten contributor categories, no definitions published as of 22 September) was compared with the stage × layer frame and left unadopted; a one-line note distinguishes RSI Bench from the audited RSI-Exam.