Per-benchmark coverage score — audit note (23 September 2026)

The overview's lead figure scores each primary benchmark by the number of distinct stage × layer cells, out of 35, in which at least one of its task verifiers executes the real LLM workload (a training run, a finetune, or a kernel on the declared hardware). This is the measured evidence class of the 5 September method note. A lighter extension adds cells reached only by bounded (unit or artifact checks in a narrower harness) or adjacent (a fixed dependency, a simulation, or a nearby non-LLM task) evidence. Family entries count once. The score counts cells, never tasks, so a task that reaches two cells is not double-counted and per-cell task counts are never summed.

Scores

Benchmark Entries Cells with any evidence Cells measured Score
RE-Bench 7 12 10 29%
MLS-Bench 14 13 7 20%
Terminal-Bench 4.0 8 12 5 14%
RSI-Exam 8 10 5 14%
PostTrainBench 1 (×28) 4 1 3%
InferenceBench 1 (×4) 4 1 3%
Terminal-Bench 2.1 8 6 0 0%
DeepSWE 3 3 0 0%
WeirdML v3 1 (×11) 2 0 0%
EdgeBench 1 (×51) 2 0 0%
Union of the ten 52 25 12 34%

Why four benchmarks score zero (verifiers re-read on 23 September)

The 35 cells and the 10 that no task reaches

A cell is one life-cycle stage (data; pretrain & midtrain; post-train; agent post-training; serve & deployment) crossed with one infrastructure layer (model & algorithm; data pipeline & evaluation; training & inference engines; parallelism & communication; kernels & compilers; cluster orchestration & sandboxing; memory & state management). Each layer carries a contract stating what a test there has to do; a task lands in a cell when its verifier reaches that stage and exercises code at that layer. Fifteen of the 52 entries sit in two stages.

Empty cell What documented production systems do there
Data × engines Tokenization and dataloader throughput (Dolma, litdata, Megatron loaders)
Data × state Loader checkpoint and exactly-once resume (worker ownership in OLMo and SmolLM3 tooling)
Data × kernels GPU dedup and tokenization kernels (NeMo Curator style pipelines)
Data × cluster Corpus builds on Slurm or Ray clusters (Dolma 3, Datatrove)
Pretrain × cluster Node health checks and cordoning; 419 unexpected interruptions in Llama 3; OLMo 2 host failures
Post-training × data pipeline & evaluation Sample mixing and mask semantics (MiMo Sample Mixer; tool-token masking in slime)
Post-training × cluster Trainer and sampler job placement (colocated verl; disaggregated GLM-5)
Agent post-training × distributed Weight publication to rollout replicas (checkpoint-engine; slime delta sync)
Agent post-training × kernels Trainer/sampler numeric parity (MiMo MXFP4 QDQ; R3 routing replay)
Agent post-training × cluster Sandbox fleets (MiMo Harness Pool; Mercor MCP fleets; SkyRL on Modal)

Sources for the production examples are the pipeline, RL-loop and agentic tables of the overview and deep-dive pages.