The overview's lead figure scores each primary benchmark by the number of distinct stage × layer cells, out of 35, in which at least one of its task verifiers executes the real LLM workload (a training run, a finetune, or a kernel on the declared hardware). This is the measured evidence class of the 5 September method note. A lighter extension adds cells reached only by bounded (unit or artifact checks in a narrower harness) or adjacent (a fixed dependency, a simulation, or a nearby non-LLM task) evidence. Family entries count once. The score counts cells, never tasks, so a task that reaches two cells is not double-counted and per-cell task counts are never summed.
| Benchmark | Entries | Cells with any evidence | Cells measured | Score |
|---|---|---|---|---|
| RE-Bench | 7 | 12 | 10 | 29% |
| MLS-Bench | 14 | 13 | 7 | 20% |
| Terminal-Bench 4.0 | 8 | 12 | 5 | 14% |
| RSI-Exam | 8 | 10 | 5 | 14% |
| PostTrainBench | 1 (×28) | 4 | 1 | 3% |
| InferenceBench | 1 (×4) | 4 | 1 | 3% |
| Terminal-Bench 2.1 | 8 | 6 | 0 | 0% |
| DeepSWE | 3 | 3 | 0 | 0% |
| WeirdML v3 | 1 (×11) | 2 | 0 | 0% |
| EdgeBench | 1 (×51) | 2 | 0 | 0% |
| Union of the ten | 52 | 25 | 12 | 34% |
7131e43): every task declares gpus = 0; GPU tasks arrived in 3.0. The tensor-parallel task starts gloo process groups on CPU at world sizes 1, 2 and 4 and compares a 2 × 64 input against nn.Linear with allclose. The pipeline task builds a 4-layer LLaMA (hidden 512) on CPU and compares hooked activations with a reference. gpt2-codegolf compiles the C file, generates a few tokens and string-matches. hf-model-inference loads a Hugging Face sentiment model and checks six POST responses. The batching task is scored by an analytical cost model. Five cells are bounded, one adjacent, none measured.0b9fabb): all 113 tasks declare gpus = 0. The three infrastructure-adjacent tasks run go test -run TestCoalesceValues (Helm), the Arcane drift-detection service and handler tests, and pytest test_coalesce.py (LangChain). Real deployment tooling, but unit tests, and adjacent to LLM infrastructure.A cell is one life-cycle stage (data; pretrain & midtrain; post-train; agent post-training; serve & deployment) crossed with one infrastructure layer (model & algorithm; data pipeline & evaluation; training & inference engines; parallelism & communication; kernels & compilers; cluster orchestration & sandboxing; memory & state management). Each layer carries a contract stating what a test there has to do; a task lands in a cell when its verifier reaches that stage and exercises code at that layer. Fifteen of the 52 entries sit in two stages.
| Empty cell | What documented production systems do there |
|---|---|
| Data × engines | Tokenization and dataloader throughput (Dolma, litdata, Megatron loaders) |
| Data × state | Loader checkpoint and exactly-once resume (worker ownership in OLMo and SmolLM3 tooling) |
| Data × kernels | GPU dedup and tokenization kernels (NeMo Curator style pipelines) |
| Data × cluster | Corpus builds on Slurm or Ray clusters (Dolma 3, Datatrove) |
| Pretrain × cluster | Node health checks and cordoning; 419 unexpected interruptions in Llama 3; OLMo 2 host failures |
| Post-training × data pipeline & evaluation | Sample mixing and mask semantics (MiMo Sample Mixer; tool-token masking in slime) |
| Post-training × cluster | Trainer and sampler job placement (colocated verl; disaggregated GLM-5) |
| Agent post-training × distributed | Weight publication to rollout replicas (checkpoint-engine; slime delta sync) |
| Agent post-training × kernels | Trainer/sampler numeric parity (MiMo MXFP4 QDQ; R3 routing replay) |
| Agent post-training × cluster | Sandbox fleets (MiMo Harness Pool; Mercor MCP fleets; SkyRL on Modal) |
Sources for the production examples are the pipeline, RL-loop and agentic tables of the overview and deep-dive pages.