What it is. AutoLab (Zhangchen Xu et al., "AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?", arXiv 2606.05080; website autolab.moe) is a public benchmark of 36 optimization tasks in Harbor format, released under Apache-2.0. Its README calls it a live benchmark, gives the current version as v1.1.26.05.10 and takes new tasks through a contribution portal. The README sorts the tasks into model development (7), system optimization (15), puzzle and challenge (10) and CUDA (4). Each task hands the agent a working but unoptimized program, a time budget of 2 to 12 hours and a metric, and its verifier writes a continuous reward between 0 and 1, anchored to a baseline and a reference score. 11 tasks declare one GPU (8 an H100, 3 an L40S), 25 declare none, and all 36 set allow_internet = false. AutoLab is the twelfth benchmark of the audit. It was read on 1 October 2026 at commit 4127da3dde8449be61a1cf9859473b9fbbd51751, pushed on 30 August 2026 and still the latest push on the day of reading.
How we read it. All 36 task folders were read at the pinned commit through the GitHub API: the manifest, the instruction and the verifier (tests/test.sh and its helper scripts) of every task, by one reader per task. A second agent per task, the skeptic, re-read the same files and tried to refute each class, cell and depth; one entry changed (scaling_law lost its Pretrain & midtrain × Memory & state management cell, because the starter code already writes the checkpoint). Two judges then ruled independently on the 13 tasks whose class or cells could bear on a claim of the post, one arguing from the post's precedents and one arguing against the post. Where the judges agree, their ruling is recorded. Where they disagree, the stricter grade is recorded. Nothing from the repository was cloned, built or run, and binaries and data files were not fetched. All entries carry the tag agent-reported. The numbers below are computed from the audit's data files by a script that first reproduces the figure for the eleven earlier benchmarks (SWE-Serve at 14%, RE-Bench at 30%, a union of 12 measured cells and 25 with any evidence).
Depth words. Each cell is graded measured (the verifier itself runs a real LLM workload), bounded (unit or artifact tests) or adjacent (related or simulated evidence). The figure in the post labels the same three depths executed, bounded and related.
data_select_ifeval, scaling_law, grpo_multisource, multilingual_ocr, llm_online_serving and flash_attention. The first five are five of AutoLab's seven model-development tasks; the other two model-development tasks train an image diffusion model and a video model and are class A.grpo_multisource and multilingual_ocr train vision-language models, which the census rule counts as language models, as it does for the existing RSI-Exam entry clevr_cogent_grpo_qwen2vl.data_select_ifeval. The twelve benchmarks together run a real LLM workload in 13 of the 35 cells (12 before); cells with any evidence stay at 25.flash_attention is a CPU kernel.| Layer | Data | Pretrain & midtrain | Post-training | Agent post-training | Serve & deployment |
|---|---|---|---|---|---|
| Model & algorithm | measured | measured | measured | · | · |
| Data pipeline & evaluation | · | · | · | · | · |
| Training & inference engines | · | adjacent | adjacent | · | measured |
| Parallelism & communication | · | · | · | · | · |
| Kernels & compilers | · | · | · | · | bounded |
| Cluster orchestration & sandboxing | · | · | · | · | · |
| Memory & state management | · | · | · | · | adjacent |
Counts. The twelve benchmarks hold 626 public task environments (590 before AutoLab), of which 156 touch the LLM lifecycle (150 before). The share stays at 25% (25.4% before, 24.9% now). Without SWE-Serve the share is 103 of 573 (18%); before AutoLab it was 97 of 537 (18%). Classes A and O go from 161 and 279 to 176 and 294. The figure holds 133 mapped entries (127 before).
Union. The eleven earlier benchmarks: 12 measured cells, 25 with any evidence, 10 empty. The twelve: 13 measured, 25 with any evidence, 10 empty. The only cell that changes is Data × Model & algorithm, adjacent before (PostTrainBench and WeirdML) and measured through data_select_ifeval. Every other AutoLab cell was already reached at the same or a deeper kind. The union's weighted score goes from 16.75 to 17.5 of 35.
Statement by statement.
| Quantity | Eleven benchmarks | Twelve, with AutoLab |
|---|---|---|
| Cells in which the benchmarks together run a real LLM workload | 12 of 35 | 13 of 35 |
| Measured cells in the Data stage | none | one, Data × Model & algorithm; the other six Data cells have no measured task environment |
| Measured cells in the Agent post-training stage | none | none; AutoLab has no Agent post-training cell at any depth |
| Measured cells in the cluster layer | none | none; AutoLab has no cluster cell at any depth |
| Most measured cells in a single benchmark | 10 (RE-Bench) | 10 (RE-Bench); AutoLab has 4 |
The two judges agree on every cell behind this table: both rule Data × Model & algorithm measured for data_select_ifeval, both give grpo_multisource only an adjacent Data cell, and both find no Agent post-training or cluster cell. They differ on two adjacent kernel cells and on the class (A or O) of two puzzle tasks, which move no measured cell and no count the post prints.
What the measured Data cell rests on. In data_select_ifeval the agent may edit one file, the selection script. The verifier reruns that script, then a fixed LoRA fine-tune of Qwen2.5-3B-Instruct on the selected samples, then full IFEval through vLLM, on one H100, and nothing is self-reported. That is the post's rule for measured, and it is how the earlier benchmarks are graded when a single edited component is scored through a rerun training job (the MLS-Bench pretraining tasks). The evidence is narrow. It is one task and one cell. The measurement is one fine-tune of the agent's selection, from one seed, read through one accuracy number on about 541 prompts (the verifier also retrains a random-selection baseline in the same run). The data step itself is a single-process filter over a 50,000-sample JSON file. No data engine, sharding, loader state, preprocessing kernel or cluster job is touched, in AutoLab or anywhere else in the audited set.
The strongest counterargument. PostTrainBench's data cell is adjacent with the note "scored only through the final model", and that phrase is also true of data_select_ifeval. On the reading that a Data cell cannot be measured through a post-training run, the cell is adjacent (or bounded), the union stays at 12 measured cells and the Data stage keeps no measured task environment. Both judges rejected that reading, for two reasons. First, PostTrainBench's verifier never runs the data step and data is one choice among many there, while here the selection script is the only editable component and the verifier reruns it. Second, that reading does not stop at AutoLab: the same rule has to be re-applied to the earlier benchmarks (MLS-Bench's llm-qat-algorithm is measured at a serve cell through a fine-tune). The audit records measured.
All rulings. Thirteen tasks went to the two judges. "Precedent" is the judge arguing from the post's existing entries; "critic" is the judge arguing against the post.
| Task | Question | Precedent judge | Critic judge | Recorded |
|---|---|---|---|---|
data_select_ifeval |
Data × Model & algorithm: is it Data-stage work, and how deep? | measured | measured | measured |
data_select_ifeval |
Post-training × Data pipeline & evaluation, as "sample mixing"? | no cell | no cell | no cell |
data_select_ifeval |
Post-training × Model & algorithm? | no cell (the recipe is frozen) | no cell (adjacent is defensible) | no cell |
data_select_ifeval |
Post-training × Training & inference engines, for the fixed scaffold? | adjacent | not contested | adjacent |
data_select_ifeval |
Any other Data cell, or the cluster layer? | no cell | no cell | no cell |
grpo_multisource |
Data × Model & algorithm, for mixing three supplied sources? | adjacent | adjacent | adjacent |
grpo_multisource |
Agent post-training? | not ruled | no cell | no cell |
agent_tool_routing |
Class L with an Agent post-training cell? | class A, no cell | class A, no cell | class A |
safety_router |
Class L with a serve cell? | class A, no cell | class A, no cell | class A |
flash_attention |
Class, and depth at Serve & deployment × Kernels & compilers? | L, bounded | L, bounded | L, bounded |
flash_attention |
Carry the kernel into Pretrain & midtrain and Post-training as adjacent? | no | yes | no (stricter grade) |
multilingual_ocr |
Class L or vision ML? | L, Post-training × Model & algorithm measured | the same | L, measured |
adaptive_compression |
Class | O | O | O |
bm25_search_go |
Class | A | A | A |
concurrent_kv_wal |
Class | A | A | A |
sha256_throughput |
Class | O | O | O |
sstable_compaction_rs |
Class | A | A | A |
toy_isa_opt |
Class | A | O | O (stricter grade) |
vliw_scheduler |
Class | A | O | O (stricter grade) |
Where the judges disagreed.
flash_attention, kernel cells in the training stages. The critic judge keeps the first reader's adjacent cells in Pretrain & midtrain and Post-training, following the shared-primitive precedent for attention kernels. The precedent judge drops them, because that precedent is worded for GPU kernels and this is a forward-only function on one CPU thread. The audit drops them: flash_attention keeps only its bounded Serve & deployment × Kernels & compilers cell. Both dropped cells are already measured in the union, so only AutoLab's own score depends on the choice: 15% as recorded, 16% (5.75 of 35) under the critic judge's ruling.toy_isa_opt and vliw_scheduler, class A or O. The precedent judge follows EdgeBench's vliw_kernel_optimization, which the census classed A. The critic judge keeps O: the fully read Terminal-Bench 2.1 simulator puzzles are O, AutoLab files both tasks under puzzle and challenge, and the EdgeBench verdict was screened from a README table. The audit records O, which matches its rule that generic puzzles are class O. No coverage number depends on it.Qualifications recorded with the rulings.
grpo_multisource. Both judges note that dropping its adjacent Data cell is as consistent with the RSI-Exam precedents as keeping it. The cell is kept. data_select_ifeval dominates it in AutoLab's row either way.flash_attention. Both judges record that its correctness gate checks only an output sum (see the verifier gaps below) and that a strict reader could lower the cell to adjacent or class the task A. The audit records class L and bounded.multilingual_ocr. Both judges class it L, following the census rule that includes vision-language models: the model is a vision-language model whose text decoder is fine-tuned. A reader who holds OCR to be vision ML classes it A. Its two cells are not new to the union.agent_tool_routing and safety_router. Both judges record that either task could be argued into class L with adjacent cells only, never bounded or measured. The audit records class A for both.Not sent to the judges. The other 23 tasks, among them scaling_law and llm_online_serving with their cells, rest on one reader and one skeptic. None of their cells is in the Data stage, the Agent post-training stage or the cluster layer. The weakest of those cells is llm_online_serving's adjacent Serve & deployment × Kernels & compilers cell: CUDA-graph capture is in the editable file, but a patch shipped with the task disables it.
In the style of section 04 of the post:
flash_attention), five declare one (three an H100, two an L40S), none declares two or more.llm_online_serving). The longest agent budget is 12 hours (scaling_law); the longest verifier budget is 4 hours (data_select_ifeval).gpus is never above 1, and all seven docker-compose.yaml files define a single service with one NVIDIA device.data_select_ifeval. The hash guard reads its expected values from /orig inside the agent's image and runs once, before train.sh executes the agent's script, so that script could rewrite the frozen config afterwards. Nothing checks that selected samples come from the pool. IFEval is cached in the image, so the test prompts are readable in the sandbox. One seed on about 541 prompts gives a standard error near 0.021 against a 0.10 anchor range. The reference solution is reported at about 0.42 against a 0.48 anchor. A failing selection script aborts the verifier before any reward file is written.flash_attention. The correctness gate compares only the sum of all outputs. The reader, the skeptic and both judges each recomputed it independently: replacing attention with "every row is the column mean of V" moves the sum by 3.3e-5 at the small size (gate 1e-2) and 9.1e-6 at the timed size (gate 1e-3). An averaging shortcut passes the gate.grpo_multisource and multilingual_ocr. The scored evaluation items and their answers sit in the agent's container under /orig/data, and nothing ties the adapter to the training script.scaling_law. The test split the verifier reads sits at an agent-writable path with no hash, and the verifier unpickles the checkpoint with weights_only=False. The instruction's context-length requirement produces only a warning.llm_online_serving. There is no output-quality gate. Token counts come from the engine's own return values. The baseline engine and the benchmark script sit unhashed in the agent's container. A shipped hotfix patch disables CUDA graphs, so the reference solution's main optimization is a no-op as shipped. Each run is timed once.test.sh differ from task.toml in several tasks (fredkin_sort_network, stack_machine_golf, toy_isa_opt, vliw_scheduler, resnet_bit_flip, flux2_klein_lora).These gaps are recorded beside the grades and did not lower any depth, as with the other eleven benchmarks.
data_select_ifeval's 4-hour verifier budget is unknown..npz and other data files, and the large recorded trajectories of flux2_klein_lora were not fetched. Some long harness sources (the CUDA drivers) were read in part.grpo_multisource could not be checked from the repository.task.toml paths because flux2_klein_lora has a second copy under tests/), all 36 manifests, the seven compose files, the README, and data_select_ifeval's test.sh, train.sh and train_config.yaml.Task links point to the pinned commit. Cells are stage × layer (depth).
| Task | What the agent must do | Hardware | Verifier runs a real language model | Cells |
|---|---|---|---|---|
data_select_ifeval |
Select 5,000 fine-tuning samples from a 50k pool to raise IFEval accuracy after LoRA SFT | 1 × H100; 8 CPUs, 32 GiB RAM; agent 8 h, verifier 4 h | yes: two LoRA fine-tunes of Qwen2.5-3B-Instruct and two full IFEval runs on vLLM | Data × Model & algorithm (measured); Post-training × Training & inference engines (adjacent) |
flash_attention |
Speed up scaled dot-product attention on CPU (n = 4,096, d = 64) | CPU only; 2 CPUs, 4 GiB RAM; agent 2 h, verifier 5 min | no: CPU kernel on random tensors | Serve & deployment × Kernels & compilers (bounded) |
grpo_multisource |
Post-train Qwen2.5-VL-7B with GRPO on multi-source visual math | 1 × L40S; 8 CPUs, 32 GiB RAM; agent 8 h, verifier 2 h | yes: Qwen2.5-VL-7B, 150 generations with the adapter and 50 on the base model; training not rerun | Post-training × Model & algorithm (measured); Post-training × Training & inference engines (adjacent); Data × Model & algorithm (adjacent) |
llm_online_serving |
Raise throughput and cut completion time of a gpt-oss-20b serving engine | 1 × H100; 8 CPUs, 64 GiB RAM; agent 2 h, verifier 2 h | yes: gpt-oss-20b under load, baseline engine and agent engine | Serve & deployment × Training & inference engines (measured); Serve & deployment × Memory & state management (adjacent); Serve & deployment × Kernels & compilers (adjacent) |
multilingual_ocr |
Fine-tune DeepSeek-OCR 3B with LoRA for Persian and Bengali OCR | 1 × L40S; 8 CPUs, 32 GiB RAM; agent 8 h, verifier 2 h | yes: DeepSeek-OCR 3B with the adapter, 400 generations; training not rerun | Post-training × Model & algorithm (measured); Post-training × Training & inference engines (adjacent) |
scaling_law |
Pretrain a GPT from scratch on WikiText-103 to minimize test perplexity | 1 × H100; 8 CPUs, 32 GiB RAM; agent 12 h, verifier 2 h | yes: the agent-trained LitGPT model over the WikiText-103 test split; training not rerun | Pretrain & midtrain × Model & algorithm (measured); Pretrain & midtrain × Training & inference engines (adjacent) |
The six mapped entries, in the audit's task schema, are in data/tasks_v5_autolab.json. All 36 verdicts are also in the task census.
Class L touches the LLM lifecycle and is mapped. Class A is machine learning on non-language models, or systems infrastructure outside an LLM role. Class O is everything else. The split between A and O for systems-flavoured tasks is a judgement call and enters no coverage number. Reasons are paraphrased; no task text is reproduced. Task links point to the pinned commit.
| Task | AutoLab category | GPU declared | Class | Why |
|---|---|---|---|---|
data_select_ifeval |
Model development | 1 × H100 | L | The verifier reruns the agent's data-selection script, runs a fixed LoRA fine-tune of Qwen2.5-3B-Instruct on the selected samples and evaluates the adapter on full IFEval through lm-eval's vLLM backend on one H100; choosing the fine-tuning data mixture for a language model is lifecycle data work. |
flash_attention |
System optimization | none | L | The verifier builds and times a single-head scaled dot-product attention function at n = 4,096, d = 64 on CPU behind an output-sum check; attention is a primitive the precedents count as an LLM kernel, but this is a CPU microbenchmark on random tensors, not a GPU kernel or a model run. |
grpo_multisource |
Model development | 1 × L40S | L | The verifier loads Qwen2.5-VL-7B with the agent's LoRA adapter on one L40S and measures held-out MathVista accuracy behind a VQA retention gate; RL post-training of a vision-language model, which the census rule counts as a language model. |
llm_online_serving |
Model development | 1 × H100 | L | The verifier loads gpt-oss-20b in the agent-edited SimpleLLM engine on one H100 and replays 96 Poisson-arrival generation requests twice (original engine, then the agent's), scoring token throughput and mean completion time; that is serving a language model. |
multilingual_ocr |
Model development | 1 × L40S | L | The verifier loads DeepSeek-OCR 3B plus the agent's LoRA adapter on one L40S and greedy-decodes text for 400 held-out images, scoring character error rate; the model trained is a vision-language model whose text decoder is fine-tuned, which the census rule counts as a language model. |
scaling_law |
Model development | 1 × H100 | L | The verifier loads the agent's from-scratch LitGPT decoder checkpoint on one H100 and computes perplexity over the WikiText-103 test split; the held-out quality of a pretrained language model is what is scored. |
agent_tool_routing |
System optimization | none | A | The verifier feeds template-generated tool schemas and queries to a pure-Python lexical retriever, gates on MRR@10 and Recall@10 and scores runtime; search infrastructure with an agent framing, and no language model, embedding model or agent is run. |
bm25_search_go |
System optimization | none | A | The verifier builds a Go BM25 engine, checks top-10 results on a small synthetic corpus and times queries over 4,000 synthetic documents; search-engine infrastructure with no language model. |
concurrent_kv_wal |
System optimization | none | A | The verifier builds a Go key-value store, diffs its output on a fixed workload against a reference and times a four-goroutine workload; storage-engine concurrency work with no language-model role. |
fft_rust |
System optimization | none | A | The verifier builds a Rust crate, checks a real-signal DFT against a pure-Python reference at small sizes and times it at n = 32,768 on CPU; a generic numeric kernel with no language-model role. |
flux2_klein_lora |
Model development | 1 × L40S | A | The verifier generates images with a FLUX.2 klein 9B diffusion transformer plus the agent's LoRA and scores CLIP and DINO similarity on one L40S; the trained model is an image generator and the Qwen3 text encoder is a frozen input. |
gaussian_blur |
System optimization | none | A | The verifier compares a blurred 256 × 256 image with a Python reference and times five 17 × 17 passes over a 4,096 × 4,096 image on CPU; an image-processing numeric kernel with no language-model role. |
huffman_canonical_decode_cuda |
CUDA | 1 × H100 | A | The verifier builds a CUDA kernel, checks byte-exact decoding of 2,048 synthetic bitstreams and times it on one H100; GPU kernel work outside any LLM role. |
icp_correspondence_step_cuda |
CUDA | 1 × H100 | A | The verifier builds a CUDA kernel for one Iterative-Closest-Point correspondence step, checks it against a CPU brute-force reference and times it on one H100; point-cloud geometry with no language-model role. |
moving_mnist_world_model |
Model development | 1 × H100 | A | The verifier loads the agent's trained checkpoint and scores 10-step rollout PSNR on 1,000 freshly generated Moving MNIST clips; GPU training of a convolutional video model, not a language model. |
msm_pippenger_bls12_381_cuda |
CUDA | 1 × H100 | A | The verifier checks bit-exact elliptic-curve multi-scalar multiplication and times it with CUDA events on one H100; a GPU kernel for a cryptographic primitive that no LLM stack runs. |
ntt_butterfly_cuda |
CUDA | 1 × H100 | A | The verifier checks a bit-exact batched forward NTT over a 64-bit prime field and times it with CUDA events on one H100; a GPU kernel for zero-knowledge and lattice cryptography, not an operation LLM stacks run. |
resnet_bit_flip |
Puzzle and challenge | none | A | The verifier applies the agent's list of float32 bit flips to a frozen 77K-parameter CIFAR-10 CNN and scores the number of flips needed to push accuracy below 12%; machine-learning security on a vision model. |
safety_router |
Puzzle and challenge | none | A | The verifier reruns the agent's NumPy training script on fixed TF-IDF/SVD feature vectors, gates a two-layer MLP on a private split and scores its parameter count; classic NLP classification with no transformer built, trained, evaluated or served. |
smallest_game_player |
Puzzle and challenge | none | A | The verifier trains the agent's model on board states and scores hidden-split move accuracy and learned-parameter count; the starter is a tiny transformer over board tokens, not a language model. |
sstable_compaction_rs |
System optimization | none | A | The verifier builds the Rust crate, checks compaction checksums and live-entry counts against constants and times a single-threaded merge of synthetic sorted tables; storage-engine work with no LLM role. |
adaptive_compression |
Puzzle and challenge | none | O | The verifier streams nine hidden synthetic byte sequences through the agent's Python predictor on CPU and scores byte-weighted cross-entropy; a compression puzzle with no language model and no natural-language data. |
adversarial_splay |
Puzzle and challenge | none | O | The verifier runs a fixed Python splay tree over the agent's 4,096-key access list and counts rotations; a data-structure puzzle with no machine learning or language-model role. |
aes128_ctr |
System optimization | none | O | The verifier builds the agent's C routine, checks it against a pure-Python AES reference and times encryption of 256 MiB on CPU; cryptography throughput, not a primitive the precedents treat as an LLM kernel. |
bvh_raytracer |
System optimization | none | O | The verifier builds the agent's C++ ray-triangle code, compares a render checksum with the baseline's and times five renders; computer graphics with no machine learning or language-model role. |
discover_sorting |
Puzzle and challenge | none | O | The verifier calls the agent's generator, checks the comparator network on all 65,536 binary inputs and scores the comparator count; a combinatorial puzzle. |
fredkin_sort_network |
Puzzle and challenge | none | O | The verifier simulates a reversible circuit over all 256 inputs and scores its gate count; a logic puzzle. |
hash_join |
System optimization | none | O | The verifier checks match count and checksum of an integer-key equi-join and times it at 20,000 × 5,000,000 rows on CPU; a generic database algorithm task. |
levenshtein_distance |
System optimization | none | O | The verifier builds the agent's C function, checks it against reference edit distances and times one million string pairs on CPU; a generic string-algorithm speed task with no model. |
radix_sort |
System optimization | none | O | The verifier builds the agent's C sort, checks small arrays against Python's sorted() and times 50 million uint32 values on CPU; a generic algorithm speed task with no model. |
regex_engine |
System optimization | none | O | The verifier builds the agent's Rust regex matcher, compares match counts with Python's re on 400 haystacks and times the pattern set over 100,000 haystacks on CPU; an automata speed task with no model. |
sha256_throughput |
System optimization | none | O | The verifier checks SHA-256 digests against hashlib and times hashing a 512 MiB buffer on CPU; a cryptographic primitive speedup, not a primitive the precedents treat as an LLM kernel. |
stack_machine_golf |
Puzzle and challenge | none | O | The verifier runs the agent's stack-machine program in a supplied C simulator on four seeds and scores the executed instruction count of a 256-element integer dot product; a programming puzzle. |
toy_isa_opt |
Puzzle and challenge | none | O | The verifier runs the agent's assembly in a supplied C pipeline simulator on four seeds and scores simulated cycles for a 512-element integer dot product; a toy-ISA puzzle, not a kernel on real hardware. The precedent judge classed it A after EdgeBench's VLIW kernel task; the stricter grade, O, is recorded. |
vliw_scheduler |
Puzzle and challenge | none | O | The verifier compiles the agent's C scheduler, checks that 3,000 synthetic ops are packed into three-slot bundles without hazards and scores the bundle count; a compiler-scheduling puzzle on a simulated machine. The precedent judge classed it A after EdgeBench's VLIW kernel task; the stricter grade, O, is recorded. |
z_order_range_scan |
System optimization | none | O | The verifier builds the Rust crate, compares range-count results with a slow reference on seeded cases and times the queries; a spatial data-structure speedup with no machine learning or LLM role. |