# Per-benchmark coverage score — audit note (23 September 2026) The overview's lead figure scores each primary benchmark by the number of distinct stage × layer cells, out of 35, in which at least one of its task verifiers **executes the real LLM workload** (a training run, a finetune, or a kernel on the declared hardware). This is the `measured` evidence class of the 5 September method note. A lighter extension adds cells reached only by `bounded` (unit or artifact checks in a narrower harness) or `adjacent` (a fixed dependency, a simulation, or a nearby non-LLM task) evidence. Family entries count once. The score counts cells, never tasks, so a task that reaches two cells is not double-counted and per-cell task counts are never summed. ## Scores | Benchmark | Entries | Cells with any evidence | Cells measured | Score | |---|---|---|---|---| | RE-Bench | 7 | 12 | 10 | 29% | | MLS-Bench | 14 | 13 | 7 | 20% | | Terminal-Bench 4.0 | 8 | 12 | 5 | 14% | | RSI-Exam | 8 | 10 | 5 | 14% | | PostTrainBench | 1 (×28) | 4 | 1 | 3% | | InferenceBench | 1 (×4) | 4 | 1 | 3% | | Terminal-Bench 2.1 | 8 | 6 | 0 | 0% | | DeepSWE | 3 | 3 | 0 | 0% | | WeirdML v3 | 1 (×11) | 2 | 0 | 0% | | EdgeBench | 1 (×51) | 2 | 0 | 0% | | **Union of the ten** | 52 | 25 | 12 | **34%** | ## Why four benchmarks score zero (verifiers re-read on 23 September) - **Terminal-Bench 2.1** (commit `7131e43`): every task declares `gpus = 0`; GPU tasks arrived in 3.0. The tensor-parallel task starts gloo process groups on CPU at world sizes 1, 2 and 4 and compares a 2 × 64 input against `nn.Linear` with `allclose`. The pipeline task builds a 4-layer LLaMA (hidden 512) on CPU and compares hooked activations with a reference. `gpt2-codegolf` compiles the C file, generates a few tokens and string-matches. `hf-model-inference` loads a Hugging Face sentiment model and checks six POST responses. The batching task is scored by an analytical cost model. Five cells are `bounded`, one `adjacent`, none `measured`. - **DeepSWE v1.1** (commit `0b9fabb`): all 113 tasks declare `gpus = 0`. The three infrastructure-adjacent tasks run `go test -run TestCoalesceValues` (Helm), the Arcane drift-detection service and handler tests, and `pytest test_coalesce.py` (LangChain). Real deployment tooling, but unit tests, and adjacent to LLM infrastructure. - **WeirdML v3**: a real GPU sandbox (four CPUs and one GPU per the task page of 16 September 2026), but the tasks are point-cloud classification, ship detection, radio-telescope calibration and the 17-task "Bonanza". Real training of non-LLM models; no LLM life-cycle stage is touched. - **EdgeBench**: the 51 public tasks in the README table are Scientific & ML (4: bipedal-walker RL, gravity inversion, GNN node classification, source inversion), Systems & SE (12: ANN vector search, VLIW kernel optimization, ffmpeg swscale, git in Zig, and others), Optimization (15), Knowledge (4), Formal (8) and Games (8). None trains, serves or checkpoints an LLM or writes a GPU kernel. ## The 35 cells and the 10 that no task reaches A cell is one life-cycle stage (data; pretrain & midtrain; post-train; agent post-training; serve & deployment) crossed with one infrastructure layer (model & algorithm; data pipeline & evaluation; training & inference engines; parallelism & communication; kernels & compilers; cluster orchestration & sandboxing; memory & state management). Each layer carries a contract stating what a test there has to do; a task lands in a cell when its verifier reaches that stage and exercises code at that layer. Fifteen of the 52 entries sit in two stages. | Empty cell | What documented production systems do there | |---|---| | Data × engines | Tokenization and dataloader throughput (Dolma, litdata, Megatron loaders) | | Data × state | Loader checkpoint and exactly-once resume (worker ownership in OLMo and SmolLM3 tooling) | | Data × kernels | GPU dedup and tokenization kernels (NeMo Curator style pipelines) | | Data × cluster | Corpus builds on Slurm or Ray clusters (Dolma 3, Datatrove) | | Pretrain × cluster | Node health checks and cordoning; 419 unexpected interruptions in Llama 3; OLMo 2 host failures | | Post-training × data pipeline & evaluation | Sample mixing and mask semantics (MiMo Sample Mixer; tool-token masking in slime) | | Post-training × cluster | Trainer and sampler job placement (colocated verl; disaggregated GLM-5) | | Agent post-training × distributed | Weight publication to rollout replicas (checkpoint-engine; slime delta sync) | | Agent post-training × kernels | Trainer/sampler numeric parity (MiMo MXFP4 QDQ; R3 routing replay) | | Agent post-training × cluster | Sandbox fleets (MiMo Harness Pool; Mercor MCP fleets; SkyRL on Modal) | Sources for the production examples are the pipeline, RL-loop and agentic tables of the overview and deep-dive pages.