This note backs the "life cycles" figure, the code-by-stage band and the RL-loop anatomy. It merges the companion notes written earlier the same day (../../llm_infra_review/production_lifecycle_sources_v2.md, rl_pipeline_sources_v2.md; tag note-2026-09-21), three delegated sweeps (tag agent-reported), and eight primary pages re-read in this session (tag spot-checked): the slime README, GLM-4.5 §3.5, the Olmo 3 blog, the checkpoint-engine README, the Terminal-Bench 2.1 news post and task packages, and the PostTrainBench and InferenceBench READMEs. Earlier workspace notes are prior-note. Numbers are in data/pipelines.json and data/rl_loop.json; re-read agent-reported entries before quoting them in a paper.
| Source | Data | Pretrain | Midtrain | Post-train | Agentic RL | Serve |
|---|---|---|---|---|---|---|
| Llama 3 (report) | curation, dedup, mixture | 16K H100, 4D parallelism | long-context continuation, annealing | rejection sampling, SFT, DPO rounds | — | inference optimization, FP8 |
| DeepSeek-V3 / R1 (report; serving code) | 14.8T tokens | 2,048 H800, DualPipe, FP8 | 32K → 128K (YaRN) | cold start → reasoning RL → rejection sampling + SFT → RL for all scenarios | — | prefill unit 4 nodes, decode unit 40 nodes; DeepEP/DeepGEMM/EPLB open |
| Qwen3 (report) | > 30T tokens | General S1, Reasoning S2 | Long Context stage | long-CoT cold start → reasoning RL → thinking-mode fusion → general RL | Qwen3.5: disaggregated async RL (blog only) | — |
| Kimi K2 / K3 (report; two components open) | 15.5T tokens | MuonClip, zero loss spike | 4K → 128K; K3 to 1M | SFT, joint RL; checkpoint-engine (1T update < 30 s) | agentic data synthesis, partial rollouts; AgentENV microVMs | K3 §5.4 |
| GLM-4.5 (report; RL on open slime) | corpus | 355B-A32B | repo-level code · synthetic reasoning · long-context & agent | expert training → unified training; SFT, reasoning/agentic/general RL on slime | disaggregated async pools, trajectory pool, Docker runtime | SGLang commands |
| GLM-5 (report; slime) | 27T tokens | 744B-A40B | 4K → 128K → 200K | SFT → reasoning RL → general RL → cross-stage distillation | fully decoupled, > 1k concurrent rollouts, heartbeat fault tolerance | PD disaggregation in the loop |
| Llama 4 (blog) | — | > 30T tokens, FP8, 32K GPUs | mid-training for 10M context | lightweight SFT → online RL → lightweight DPO | fully async online RL framework | — |
| DeepSeek-V4 (report) | — | 32–33T tokens, FP4 experts | 1M context | domain specialists → on-policy distillation | preemptible generation service (token WAL); DSec sandboxes | on-disk KV cache |
| MiniMax-M2 (report) | — | 29.2T tokens | 8K → 32K → 192K | SFT, CISPO RL | Forge: gateway, data pool, windowed FIFO | — |
| Nemotron 3 Super (report + open recipe) | data-prep (subset) | Megatron-Bridge, 25T tokens | LC-phase CPT at 1M on GB200 | SFT → RLVR/SWE-RL/RLHF → MTP healing (NeMo-RL) | SWE-RL, NeMo-Gym, 4,000 env instances per batch | quantization, evaluator |
| OLMo 3 (fully open) | Dolma 3 tools | OLMo-core, 5.9T, up to 1,024 H100 | Dolmino 100B; Longmino ~50B | Open Instruct: SFT → DPO → RLVR (OlmoRL async) | — | — |
| SmolLM3 (fully open) | datatrove | nanotron, 3 stages to 11.1T, 384 H100 | long-context 100B; reasoning mid-training 140B | TRL: SFT → APO, merging | — | — |
| Apertus (open pretraining) | pretrain-data | Megatron fork, up to 4,096 GPUs | 4K → 64K | SFT → QRPO | — | — |
| LLM360 K2 / K2-V2 (fully open) | TxT360 | 480 A100, weekly hardware failures | four mid-training stages to 512K | SFT | — | — |
| Marin (open lab) | data browser, executor DAG | Levanter on TPU, 12.7T (8B) | — | SFT; RL by 2026 | — | — |
| Arcee Trinity Large (report; RL on prime-rl) | — | 2,048 B300, modified TorchTitan | ~117B long-context | SFT at 64K; RL with prime-rl | — | — |
| Mercor 397B (open recipe) | — | — | — | — | SkyRL async, S=3, 300–550 concurrent rollouts | vLLM rollout engines |
Observations that shape the page:
The sweep decomposes every stack into six roles: a trainer (Megatron-Core in slime, AReaL, SkyRL/Mercor, NeMo-RL, ROLL and verl; FSDP2 in prime-rl, torchforge and OpenRLHF; hidden behind LoRA in Tinker), a rollout service with an OpenAI-style endpoint (vLLM or SGLang), a weight-sync path (NCCL broadcast in verl, SkyRL, Magistral, TRL and NeMo-RL; a checkpoint engine in Kimi at about 20 s per 1T model; HTTP relays in prime-rl; delta patches in slime; an RDMA store in torchforge), a queue or buffer (slime Data Buffer, verl MessageQueue and TransferQueue, NeMo-RL replay buffer, LongCat SampleQueue, MiniMax Data Pool), reward/verifier and sandbox services scaled on CPU (Prime Sandboxes > 4,000, LongCat 32,000 environments, Kimi K2.5 up to 100,000 tasks, DeepSeek DSec "hundreds of thousands"), and an orchestrator (Ray almost everywhere; Monarch in torchforge).
They differ in placement (colocated hybrid engines for reasoning RL in Kimi K2, Seed1.5-Thinking, verl and GLM-4.5's synchronous mode; disaggregated GPUs for agentic RL in GLM-5, INTELLECT-3's 16:44 split, Mercor's 12:8, NeMo-RL's async mode, RollArt), in synchrony (bounded-staleness knobs: AReaL η = 4/8, INTELLECT-3 max_off_policy_steps = 8, Mercor 4, NeMo-RL max_trajectory_age_steps = 1, verl staleness_threshold, SkyRL max_staleness_steps; in-flight updates and partial rollouts; 2–4× speedups reported), and in the train–inference mismatch fix (importance-sampling correction by default in TRL and required in NeMo-RL; FP32 LM head in MiniMax; batch-invariant kernels in Thinking Machines and TitanRL). This matches the workspace's staleness note (../../ASYNC_RL_STALENESS.md).
The five-part conceptual loop on the page (prompts & environments → rollout → score & admit → learn → publish weights, with checkpoint/recovery and placement/networking/sandbox rails) is the earlier companion note's proposal, labeled as a synthesis illustrated with slime and GLM-4.5. It is not any repository's class diagram.
Applied from the user's review of the page: prose now runs at full width; six older or thinner reports (Llama 3, DeepSeek-V3/R1, Llama 4, Apertus, LLM360 K2, Arcee Trinity) are hidden behind a toggle and dimmed when shown; Mercor stays visible. Continued pretraining, annealing and long-context extension were merged into the pretraining column because they share the objective and trainer and no audited task targets them separately. On-policy distillation joined the post-training vocabulary with four code chips (NeMo-RL's on-policy and X-token distillation recipes, TRL's GKDTrainer, Thinking Machines' Tinker post, Xiaomi's MOPD/MOPD2 paper) after re-reading each source. The Marin row now cites the Open Athena 535B launch note (August 2026 start on CoreWeave GPUs with JAX/XLA:GPU, 18T tokens, MFU 21% → 24%, preregistered loss predictions; Snowball 67B-A2B in post-training). The RL loop moved to two deep-dive pages: posttrain.html (recipes as sequences for all 18 pipelines, frameworks by sub-stage, twelve stabilizers with sources) and agentic.html (the loop with 59 implementations, ten fleets, twelve staleness entries, ten ledger rows, seven RL-Infra-Bench families). The grid no longer highlights the agentic column.
The user asked that the pages read as a review rather than as the case for a benchmark under construction. Every mention of RL-Infra-Bench on the pages was removed: the coverage figure's target box and key item, the loop drawer's "families" line, and the agentic page's blueprint-derived table, which is now "Mechanisms that only exist across devices" with three columns (mechanism, why it needs several devices, what documents it). The overview's post-training section became three blocks: the RL loop (new page loop.html, with the figure, drawer and a component table), post-training recipes, and agentic RL infrastructure. The mapping from incidents to benchmark mechanism families remains in this note's §3 and in mimo_v26_review.md §8 for the benchmark work, off the pages.
After the user's review of the published draft: the coverage figure moved to the top of the overview, the title became "AI-Infra Environment Coverage" with the question as a subtitle, every mono subhead was removed, captions and section text were cut to about half (overview prose from roughly 700 to 480 words), and the changelog became a working-copy appendix that deploy_site.py strips before publishing. Scale's RSI Bench taxonomy (ten contributor categories, no definitions published as of 22 September) was compared with the stage × layer frame and left unadopted; a one-line note distinguishes RSI Bench from the audited RSI-Exam.