# Benchmark additions, 21 September 2026 Five benchmarks were added to the ten-row coverage figure. Tags: **spot-checked** = primary page re-read in this session; **agent-reported** = delegated sweep, re-read before quoting. ## Terminal-Bench 2.1 — spot-checked Released 6 May 2026 (TB2.1 lead Kelly Buchanan). The same 89 tasks as 2.0 (the task directory at commit `7131e4375048a0e408a8fb404b5f499d726b695b`, 11 August 2026, holds exactly 89 task directories); the news post says 28 tasks were corrected, the README says 26 were modified. Fixes fall in three categories: internet-dependent images (nine tasks), too-tight resource budgets (eight), and misspecification. "After these changes, no task is unsolved." Claude Code + Opus 4.6 rose from 58.0% to 70.1%. Canonical Harbor format (`task.toml` schema 1.1). CPU only: both task.toml files read declare `gpus = 0`; GPU environments arrived in 3.0, and 4.0 shares no tasks with 2.1. Reported in Google's Gemini 3.5 Flash model card (May 2026) and Anthropic's mid-2026 system cards; labs have since moved to 4.0. Infrastructure-related tasks mapped (three re-read at the task package, five from the sweep's reading of `instruction.md`): | Task | What it tests | Mapped as | |---|---|---| | torch-tensor-parallelism | Column/row-parallel linear layers over torch.distributed; sharding, outputs and gradients at world sizes 1/2/4 on 1 CPU, 8 GB | pretrain / distributed, bounded | | torch-pipeline-parallelism | AFAB pipeline train step for a LlamaForCausalLM; hooks compare activations to a reference at world sizes 1/2 | pretrain / distributed and engines, bounded | | llm-inference-batching-scheduler | Shape-aware batching plans scored by an analytical cost model; no engine or GPU | serve / engines, related (simulated) | | gpt2-codegolf | GPT-2 inference in under 5,000 bytes of C, matched to a reference | serve / engines, bounded | | count-dataset-tokens | Exact token count with the Qwen2.5-1.5B tokenizer | data / data & eval, bounded | | reshard-c4-data | Reshard C4 under file-count and size limits with a round-trip test | data / data & eval, bounded; data / distributed, related | | pytorch-model-recovery | Rebuild a model from a state_dict, finetune the output layer, export TorchScript | post-train / state, bounded | | hf-model-inference | Flask endpoint serving an HF model | serve / engines, related | No task name matches vllm, sglang, triton, cuda, megatron, nccl, jax, kv-cache or quantization; `kv-store-grpc` is a generic key-value store. ## PostTrainBench — spot-checked AISA group (ELLIS Tübingen / MPI-IS); arXiv 2603.08640; ICML 2026 poster; MIT. 28 tasks = 4 base models × 7 benchmarks. "The agent is given access to an evaluation script and 10 hours on an H100 GPU"; it must "store your best trained model in the folder final_model"; contamination rule: "Do not use {benchmark} test data for training (neither questions, nor answers)." Apptainer containers with HTCondor; "Harbor support coming soon." Scored on held-out sets by exact match or a GPT-5-mini judge. Paper: best agent 23.2% versus 51.1% for official instruct models; leaderboard v1.1 (Sept 2026) tops out around 34%. Mapped as one family entry (×28): post-train / methods measured; data / methods, post-train / engines and post-train / state related. The training implementation is not verified, only the resulting weights, as in RSI-Exam. ## InferenceBench — spot-checked Same group; arXiv 2607.20468; Apache-2.0, v1.0.5 (Sept 2026). Four scenarios (prefill latency, decode latency, throughput, multi-objective geomean) on Mistral-7B-Instruct-v0.3, one H100 80 GB, a 2-hour budget. The agent delivers a running OpenAI-compatible server; a quality gate requires at least 0.95 × the PyTorch baseline accuracy on a 500-question MMLU-Pro subset; an agentic integrity judge screens for reward hacking and relaunches the server in a fresh container. vLLM, SGLang, TGI and TensorRT-LLM are pre-installed. Independent of ISO-Bench and FlashInfer-Bench, which it cites. A matched-budget non-agentic search (11.5×) beats every agent (up to 8.1×), so the score measures selection and tuning of existing engines. Mapped as one family entry (×4): serve / engines measured; serve / state, methods and cluster related. The workspace's earlier note in `systems_sources.md` already flagged that repository-internal patch coverage is trajectory dependent. ## WeirdML v3 — agent-reported Håvard Tveit Ihle (NDRE); funded primarily by Epoch AI, secondarily by METR and NDRE. v3 is "an agentic benchmark featuring 11 complex hand-made tasks" (exploring unfamiliar data and building ML pipelines under a 500k–50M token budget); v2 (June 2025, 19 tasks) had agents write PyTorch code that trains a model on a novel dataset in Docker on a 12-GB TITAN V. No public code repository was found. Epoch AI runs v2 and includes WeirdML in its capabilities index; no frontier model card cites it. Mapped as one family entry (×11): data / methods and data / data & eval, related. It is non-LLM machine learning; it does not touch LLM training or serving infrastructure. ## EdgeBench — agent-reported Several things carry the name. The 2026 agent benchmark is ByteDance Seed's "EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments" (arXiv 2607.05155; 134 expert-built tasks, 51 public / 83 held-out; SForge harness with a work container and a judge container per task; primary metric score@12h). "Edge" is never defined and the benchmark is not about edge devices; the on-device candidates (MLPerf Inference v6.1 Edge Agentic, edge-llm-bench) are engine performance harnesses, not agent tasks. None of the 51 public tasks targets LLM training, serving, GPU kernels or checkpoints; the systems tasks are CPU SIMD/VLIW, ANN search and compression work. Mapped as one family entry (×51): pretrain / methods (from-scratch GNN and RL training) and serve / kernels (CPU kernels), both related. ## Supplemental rows (benchmark-level, hollow rings, not counted) ISO-Bench (54 vLLM/SGLang PR tasks on H100), FlashInfer-Bench (serving kernels under the FlashInfer Trace schema), CommBench (101 multi-GPU communication tasks; reference sources withheld; see `../../COMMBENCH_DIGEST.md`), ATE-Bench (torchtitan, Megatron-LM, PithTrain tasks) and KernelBench (250 kernel problems). All from earlier workspace notes; none task-audited here.