AutoLab review, 1 October 2026

What it is. AutoLab (Zhangchen Xu et al., "AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?", arXiv 2606.05080; website autolab.moe) is a public benchmark of 36 optimization tasks in Harbor format, released under Apache-2.0. Its README calls it a live benchmark, gives the current version as v1.1.26.05.10 and takes new tasks through a contribution portal. The README sorts the tasks into model development (7), system optimization (15), puzzle and challenge (10) and CUDA (4). Each task hands the agent a working but unoptimized program, a time budget of 2 to 12 hours and a metric, and its verifier writes a continuous reward between 0 and 1, anchored to a baseline and a reference score. 11 tasks declare one GPU (8 an H100, 3 an L40S), 25 declare none, and all 36 set allow_internet = false. AutoLab is the twelfth benchmark of the audit. It was read on 1 October 2026 at commit 4127da3dde8449be61a1cf9859473b9fbbd51751, pushed on 30 August 2026 and still the latest push on the day of reading.

How we read it. All 36 task folders were read at the pinned commit through the GitHub API: the manifest, the instruction and the verifier (tests/test.sh and its helper scripts) of every task, by one reader per task. A second agent per task, the skeptic, re-read the same files and tried to refute each class, cell and depth; one entry changed (scaling_law lost its Pretrain & midtrain × Memory & state management cell, because the starter code already writes the checkpoint). Two judges then ruled independently on the 13 tasks whose class or cells could bear on a claim of the post, one arguing from the post's precedents and one arguing against the post. Where the judges agree, their ruling is recorded. Where they disagree, the stricter grade is recorded. Nothing from the repository was cloned, built or run, and binaries and data files were not fetched. All entries carry the tag agent-reported. The numbers below are computed from the audit's data files by a script that first reproduces the figure for the eleven earlier benchmarks (SWE-Serve at 14%, RE-Bench at 30%, a union of 12 measured cells and 25 with any evidence).

Depth words. Each cell is graded measured (the verifier itself runs a real LLM workload), bounded (unit or artifact tests) or adjacent (related or simulated evidence). The figure in the post labels the same three depths executed, bounded and related.

Result on the audit's surface

AutoLab's row in the coverage figure

Layer Data Pretrain & midtrain Post-training Agent post-training Serve & deployment
Model & algorithm measured measured measured · ·
Data pipeline & evaluation · · · · ·
Training & inference engines · adjacent adjacent · measured
Parallelism & communication · · · · ·
Kernels & compilers · · · · bounded
Cluster orchestration & sandboxing · · · · ·
Memory & state management · · · · adjacent

Effect on the counts and on the union

Counts. The twelve benchmarks hold 626 public task environments (590 before AutoLab), of which 156 touch the LLM lifecycle (150 before). The share stays at 25% (25.4% before, 24.9% now). Without SWE-Serve the share is 103 of 573 (18%); before AutoLab it was 97 of 537 (18%). Classes A and O go from 161 and 279 to 176 and 294. The figure holds 133 mapped entries (127 before).

Union. The eleven earlier benchmarks: 12 measured cells, 25 with any evidence, 10 empty. The twelve: 13 measured, 25 with any evidence, 10 empty. The only cell that changes is Data × Model & algorithm, adjacent before (PostTrainBench and WeirdML) and measured through data_select_ifeval. Every other AutoLab cell was already reached at the same or a deeper kind. The union's weighted score goes from 16.75 to 17.5 of 35.

Statement by statement.

Quantity Eleven benchmarks Twelve, with AutoLab
Cells in which the benchmarks together run a real LLM workload 12 of 35 13 of 35
Measured cells in the Data stage none one, Data × Model & algorithm; the other six Data cells have no measured task environment
Measured cells in the Agent post-training stage none none; AutoLab has no Agent post-training cell at any depth
Measured cells in the cluster layer none none; AutoLab has no cluster cell at any depth
Most measured cells in a single benchmark 10 (RE-Bench) 10 (RE-Bench); AutoLab has 4

The two judges agree on every cell behind this table: both rule Data × Model & algorithm measured for data_select_ifeval, both give grpo_multisource only an adjacent Data cell, and both find no Agent post-training or cluster cell. They differ on two adjacent kernel cells and on the class (A or O) of two puzzle tasks, which move no measured cell and no count the post prints.

Rulings on the claim-affecting cells

What the measured Data cell rests on. In data_select_ifeval the agent may edit one file, the selection script. The verifier reruns that script, then a fixed LoRA fine-tune of Qwen2.5-3B-Instruct on the selected samples, then full IFEval through vLLM, on one H100, and nothing is self-reported. That is the post's rule for measured, and it is how the earlier benchmarks are graded when a single edited component is scored through a rerun training job (the MLS-Bench pretraining tasks). The evidence is narrow. It is one task and one cell. The measurement is one fine-tune of the agent's selection, from one seed, read through one accuracy number on about 541 prompts (the verifier also retrains a random-selection baseline in the same run). The data step itself is a single-process filter over a 50,000-sample JSON file. No data engine, sharding, loader state, preprocessing kernel or cluster job is touched, in AutoLab or anywhere else in the audited set.

The strongest counterargument. PostTrainBench's data cell is adjacent with the note "scored only through the final model", and that phrase is also true of data_select_ifeval. On the reading that a Data cell cannot be measured through a post-training run, the cell is adjacent (or bounded), the union stays at 12 measured cells and the Data stage keeps no measured task environment. Both judges rejected that reading, for two reasons. First, PostTrainBench's verifier never runs the data step and data is one choice among many there, while here the selection script is the only editable component and the verifier reruns it. Second, that reading does not stop at AutoLab: the same rule has to be re-applied to the earlier benchmarks (MLS-Bench's llm-qat-algorithm is measured at a serve cell through a fine-tune). The audit records measured.

All rulings. Thirteen tasks went to the two judges. "Precedent" is the judge arguing from the post's existing entries; "critic" is the judge arguing against the post.

Task Question Precedent judge Critic judge Recorded
data_select_ifeval Data × Model & algorithm: is it Data-stage work, and how deep? measured measured measured
data_select_ifeval Post-training × Data pipeline & evaluation, as "sample mixing"? no cell no cell no cell
data_select_ifeval Post-training × Model & algorithm? no cell (the recipe is frozen) no cell (adjacent is defensible) no cell
data_select_ifeval Post-training × Training & inference engines, for the fixed scaffold? adjacent not contested adjacent
data_select_ifeval Any other Data cell, or the cluster layer? no cell no cell no cell
grpo_multisource Data × Model & algorithm, for mixing three supplied sources? adjacent adjacent adjacent
grpo_multisource Agent post-training? not ruled no cell no cell
agent_tool_routing Class L with an Agent post-training cell? class A, no cell class A, no cell class A
safety_router Class L with a serve cell? class A, no cell class A, no cell class A
flash_attention Class, and depth at Serve & deployment × Kernels & compilers? L, bounded L, bounded L, bounded
flash_attention Carry the kernel into Pretrain & midtrain and Post-training as adjacent? no yes no (stricter grade)
multilingual_ocr Class L or vision ML? L, Post-training × Model & algorithm measured the same L, measured
adaptive_compression Class O O O
bm25_search_go Class A A A
concurrent_kv_wal Class A A A
sha256_throughput Class O O O
sstable_compaction_rs Class A A A
toy_isa_opt Class A O O (stricter grade)
vliw_scheduler Class A O O (stricter grade)

Where the judges disagreed.

Qualifications recorded with the rulings.

Not sent to the judges. The other 23 tasks, among them scaling_law and llm_online_serving with their cells, rest on one reader and one skeptic. None of their cells is in the Data stage, the Agent post-training stage or the cluster layer. The weakest of those cells is llm_online_serving's adjacent Serve & deployment × Kernels & compilers cell: CUDA-graph capture is in the editable file, but a patch shipped with the task disables it.

GPU tally

In the style of section 04 of the post:

Verifier gaps the readers flagged

These gaps are recorded beside the grades and did not lower any depth, as with the other eleven benchmarks.

What was not verified

The six lifecycle tasks

Task links point to the pinned commit. Cells are stage × layer (depth).

Task What the agent must do Hardware Verifier runs a real language model Cells
data_select_ifeval Select 5,000 fine-tuning samples from a 50k pool to raise IFEval accuracy after LoRA SFT 1 × H100; 8 CPUs, 32 GiB RAM; agent 8 h, verifier 4 h yes: two LoRA fine-tunes of Qwen2.5-3B-Instruct and two full IFEval runs on vLLM Data × Model & algorithm (measured); Post-training × Training & inference engines (adjacent)
flash_attention Speed up scaled dot-product attention on CPU (n = 4,096, d = 64) CPU only; 2 CPUs, 4 GiB RAM; agent 2 h, verifier 5 min no: CPU kernel on random tensors Serve & deployment × Kernels & compilers (bounded)
grpo_multisource Post-train Qwen2.5-VL-7B with GRPO on multi-source visual math 1 × L40S; 8 CPUs, 32 GiB RAM; agent 8 h, verifier 2 h yes: Qwen2.5-VL-7B, 150 generations with the adapter and 50 on the base model; training not rerun Post-training × Model & algorithm (measured); Post-training × Training & inference engines (adjacent); Data × Model & algorithm (adjacent)
llm_online_serving Raise throughput and cut completion time of a gpt-oss-20b serving engine 1 × H100; 8 CPUs, 64 GiB RAM; agent 2 h, verifier 2 h yes: gpt-oss-20b under load, baseline engine and agent engine Serve & deployment × Training & inference engines (measured); Serve & deployment × Memory & state management (adjacent); Serve & deployment × Kernels & compilers (adjacent)
multilingual_ocr Fine-tune DeepSeek-OCR 3B with LoRA for Persian and Bengali OCR 1 × L40S; 8 CPUs, 32 GiB RAM; agent 8 h, verifier 2 h yes: DeepSeek-OCR 3B with the adapter, 400 generations; training not rerun Post-training × Model & algorithm (measured); Post-training × Training & inference engines (adjacent)
scaling_law Pretrain a GPT from scratch on WikiText-103 to minimize test perplexity 1 × H100; 8 CPUs, 32 GiB RAM; agent 12 h, verifier 2 h yes: the agent-trained LitGPT model over the WikiText-103 test split; training not rerun Pretrain & midtrain × Model & algorithm (measured); Pretrain & midtrain × Training & inference engines (adjacent)

The six mapped entries, in the audit's task schema, are in data/tasks_v5_autolab.json. All 36 verdicts are also in the task census.

All 36 tasks

Class L touches the LLM lifecycle and is mapped. Class A is machine learning on non-language models, or systems infrastructure outside an LLM role. Class O is everything else. The split between A and O for systems-flavoured tasks is a judgement call and enters no coverage number. Reasons are paraphrased; no task text is reproduced. Task links point to the pinned commit.

Task AutoLab category GPU declared Class Why
data_select_ifeval Model development 1 × H100 L The verifier reruns the agent's data-selection script, runs a fixed LoRA fine-tune of Qwen2.5-3B-Instruct on the selected samples and evaluates the adapter on full IFEval through lm-eval's vLLM backend on one H100; choosing the fine-tuning data mixture for a language model is lifecycle data work.
flash_attention System optimization none L The verifier builds and times a single-head scaled dot-product attention function at n = 4,096, d = 64 on CPU behind an output-sum check; attention is a primitive the precedents count as an LLM kernel, but this is a CPU microbenchmark on random tensors, not a GPU kernel or a model run.
grpo_multisource Model development 1 × L40S L The verifier loads Qwen2.5-VL-7B with the agent's LoRA adapter on one L40S and measures held-out MathVista accuracy behind a VQA retention gate; RL post-training of a vision-language model, which the census rule counts as a language model.
llm_online_serving Model development 1 × H100 L The verifier loads gpt-oss-20b in the agent-edited SimpleLLM engine on one H100 and replays 96 Poisson-arrival generation requests twice (original engine, then the agent's), scoring token throughput and mean completion time; that is serving a language model.
multilingual_ocr Model development 1 × L40S L The verifier loads DeepSeek-OCR 3B plus the agent's LoRA adapter on one L40S and greedy-decodes text for 400 held-out images, scoring character error rate; the model trained is a vision-language model whose text decoder is fine-tuned, which the census rule counts as a language model.
scaling_law Model development 1 × H100 L The verifier loads the agent's from-scratch LitGPT decoder checkpoint on one H100 and computes perplexity over the WikiText-103 test split; the held-out quality of a pretrained language model is what is scored.
agent_tool_routing System optimization none A The verifier feeds template-generated tool schemas and queries to a pure-Python lexical retriever, gates on MRR@10 and Recall@10 and scores runtime; search infrastructure with an agent framing, and no language model, embedding model or agent is run.
bm25_search_go System optimization none A The verifier builds a Go BM25 engine, checks top-10 results on a small synthetic corpus and times queries over 4,000 synthetic documents; search-engine infrastructure with no language model.
concurrent_kv_wal System optimization none A The verifier builds a Go key-value store, diffs its output on a fixed workload against a reference and times a four-goroutine workload; storage-engine concurrency work with no language-model role.
fft_rust System optimization none A The verifier builds a Rust crate, checks a real-signal DFT against a pure-Python reference at small sizes and times it at n = 32,768 on CPU; a generic numeric kernel with no language-model role.
flux2_klein_lora Model development 1 × L40S A The verifier generates images with a FLUX.2 klein 9B diffusion transformer plus the agent's LoRA and scores CLIP and DINO similarity on one L40S; the trained model is an image generator and the Qwen3 text encoder is a frozen input.
gaussian_blur System optimization none A The verifier compares a blurred 256 × 256 image with a Python reference and times five 17 × 17 passes over a 4,096 × 4,096 image on CPU; an image-processing numeric kernel with no language-model role.
huffman_canonical_decode_cuda CUDA 1 × H100 A The verifier builds a CUDA kernel, checks byte-exact decoding of 2,048 synthetic bitstreams and times it on one H100; GPU kernel work outside any LLM role.
icp_correspondence_step_cuda CUDA 1 × H100 A The verifier builds a CUDA kernel for one Iterative-Closest-Point correspondence step, checks it against a CPU brute-force reference and times it on one H100; point-cloud geometry with no language-model role.
moving_mnist_world_model Model development 1 × H100 A The verifier loads the agent's trained checkpoint and scores 10-step rollout PSNR on 1,000 freshly generated Moving MNIST clips; GPU training of a convolutional video model, not a language model.
msm_pippenger_bls12_381_cuda CUDA 1 × H100 A The verifier checks bit-exact elliptic-curve multi-scalar multiplication and times it with CUDA events on one H100; a GPU kernel for a cryptographic primitive that no LLM stack runs.
ntt_butterfly_cuda CUDA 1 × H100 A The verifier checks a bit-exact batched forward NTT over a 64-bit prime field and times it with CUDA events on one H100; a GPU kernel for zero-knowledge and lattice cryptography, not an operation LLM stacks run.
resnet_bit_flip Puzzle and challenge none A The verifier applies the agent's list of float32 bit flips to a frozen 77K-parameter CIFAR-10 CNN and scores the number of flips needed to push accuracy below 12%; machine-learning security on a vision model.
safety_router Puzzle and challenge none A The verifier reruns the agent's NumPy training script on fixed TF-IDF/SVD feature vectors, gates a two-layer MLP on a private split and scores its parameter count; classic NLP classification with no transformer built, trained, evaluated or served.
smallest_game_player Puzzle and challenge none A The verifier trains the agent's model on board states and scores hidden-split move accuracy and learned-parameter count; the starter is a tiny transformer over board tokens, not a language model.
sstable_compaction_rs System optimization none A The verifier builds the Rust crate, checks compaction checksums and live-entry counts against constants and times a single-threaded merge of synthetic sorted tables; storage-engine work with no LLM role.
adaptive_compression Puzzle and challenge none O The verifier streams nine hidden synthetic byte sequences through the agent's Python predictor on CPU and scores byte-weighted cross-entropy; a compression puzzle with no language model and no natural-language data.
adversarial_splay Puzzle and challenge none O The verifier runs a fixed Python splay tree over the agent's 4,096-key access list and counts rotations; a data-structure puzzle with no machine learning or language-model role.
aes128_ctr System optimization none O The verifier builds the agent's C routine, checks it against a pure-Python AES reference and times encryption of 256 MiB on CPU; cryptography throughput, not a primitive the precedents treat as an LLM kernel.
bvh_raytracer System optimization none O The verifier builds the agent's C++ ray-triangle code, compares a render checksum with the baseline's and times five renders; computer graphics with no machine learning or language-model role.
discover_sorting Puzzle and challenge none O The verifier calls the agent's generator, checks the comparator network on all 65,536 binary inputs and scores the comparator count; a combinatorial puzzle.
fredkin_sort_network Puzzle and challenge none O The verifier simulates a reversible circuit over all 256 inputs and scores its gate count; a logic puzzle.
hash_join System optimization none O The verifier checks match count and checksum of an integer-key equi-join and times it at 20,000 × 5,000,000 rows on CPU; a generic database algorithm task.
levenshtein_distance System optimization none O The verifier builds the agent's C function, checks it against reference edit distances and times one million string pairs on CPU; a generic string-algorithm speed task with no model.
radix_sort System optimization none O The verifier builds the agent's C sort, checks small arrays against Python's sorted() and times 50 million uint32 values on CPU; a generic algorithm speed task with no model.
regex_engine System optimization none O The verifier builds the agent's Rust regex matcher, compares match counts with Python's re on 400 haystacks and times the pattern set over 100,000 haystacks on CPU; an automata speed task with no model.
sha256_throughput System optimization none O The verifier checks SHA-256 digests against hashlib and times hashing a 512 MiB buffer on CPU; a cryptographic primitive speedup, not a primitive the precedents treat as an LLM kernel.
stack_machine_golf Puzzle and challenge none O The verifier runs the agent's stack-machine program in a supplied C simulator on four seeds and scores the executed instruction count of a 256-element integer dot product; a programming puzzle.
toy_isa_opt Puzzle and challenge none O The verifier runs the agent's assembly in a supplied C pipeline simulator on four seeds and scores simulated cycles for a 512-element integer dot product; a toy-ISA puzzle, not a kernel on real hardware. The precedent judge classed it A after EdgeBench's VLIW kernel task; the stricter grade, O, is recorded.
vliw_scheduler Puzzle and challenge none O The verifier compiles the agent's C scheduler, checks that 3,000 synthetic ops are packed into three-slot bundles without hazards and scores the bundle count; a compiler-scheduling puzzle on a simulated machine. The precedent judge classed it A after EdgeBench's VLIW kernel task; the stricter grade, O, is recorded.
z_order_range_scan System optimization none O The verifier builds the Rust crate, compares range-count results with a slow reference on seeded cases and times the queries; a spatial data-structure speedup with no machine learning or LLM role.