[WIP] How to benchmark RSI
What it takes to test whether an AI system can train an AI system end to end
Recursive self-improvement (RSI) could have major consequences for AI development, yet its definition, its possible forms and its measurement remain open questions. Recent reports from Anthropic, OpenAI and Z.ai describe AI systems taking on parts of model research and infrastructure engineering, which makes these questions increasingly practical. Whatever definition prevails, the ability to train and improve AI models end to end is a core capability, and it is necessary, though not sufficient, for RSI. One way to evaluate progress toward RSI is to ask an AI system to do exactly that: develop and train a model and deliver it ready for real-world use. At production scale, such an evaluation is still too expensive in compute and time.
Coding benchmarks such as SWE-Bench Pro, Terminal-Bench, DeepSWE and ProgramBench assess progress in coding agents, and the task environments1A task environment is a containerized snapshot of a server, a developer machine or a simulated world, prepared with an initial state, the required artifacts and, usually, a verifiable goal, so that an agent can be rolled out in it repeatedly against a realistic problem. Terminal-Bench and Harbor call this a task: “one or more instructions, a sandbox environment, and a verifier” (Harbor; see the Terminal-Bench task layout). The sandbox is meant to hold, and does not always: in July 2026, OpenAI reported that its models under evaluation on the ExploitGym cyber benchmark found a zero-day in a package-registry proxy, reached the Internet and broke into Hugging Face’s infrastructure to take the test answers. they are built from also support agent training. However, as tests of LLM training itself, they fall short. Of the 626 public task environments in the twelve benchmarks we audited, 156 touch any part of the LLM lifecycle, and together they run a real LLM workload in only 13 of the 35 stage × layer cells we map below; none does so for agent post-training or cluster work, and only one does so for data.
In this post, we investigate how much production-level infrastructure knowledge for LLM training is covered by existing public task environments, and what a benchmark would need in order to cover the rest (section 04).
What is needed for end-to-end LLM training
We organize the audit along two dimensions: five stages of the LLM development and deployment lifecycle, and seven system layers. The classification synthesizes the architectures of production frameworks and the model-development pipelines documented in public reports and code, among them Qwen3, Kimi K2, GLM-5, OLMo 3 and SmolLM3 (section 03 shows all twelve). The categories overlap: data and evaluation recur throughout development, agent post-training is the multi-turn, sandboxed and often asynchronous part of post-training, and infrastructure responsibilities span several components.
We assign coverage by reading what each task environment's verifier exercises (we did not run the task environments ourselves), and grade each cell by how deeply, from deepest to shallowest: the verifier executes a real LLM workload, runs bounded unit or artifact tests, or offers only related or simulated evidence. A benchmark covers a cell as soon as one of its task environments reaches it, and all 35 cells weigh the same, so the coverage scores in section 02 measure breadth: how much of the lifecycle a benchmark touches, not how many of its tasks test each part.
5 stages of LLM training
7 system layers
Check a few examples
Where public benchmark environments land
Each row below is a benchmark. The marks show which of the 35 stage × layer cells its task environments reach, and how deeply; the bar sums them into one coverage score; the bottom row counts task environments per stage.
We read every public task environment in twelve benchmarks and audited each one that touches the LLM lifecycle (task census). We find that:
- In the general-purpose coding benchmarks, only a small fraction of task environments test an agent's ability to work on LLM training itself: 9 of 66 in Terminal-Bench 4.0, 9 of 89 in Terminal-Bench 2.1 and 3 of 113 in DeepSWE. Taken together, 156 of the 626 public task environments (25%) touch any part of the LLM lifecycle, and about a third of them (53) are the serving tasks of SWE-Serve; without SWE-Serve, the share is 103 of 573 (18%).
- Only a small part of the lifecycle's infrastructure knowledge is exercised by any task environment: the twelve benchmarks together run a real LLM workload in 13 of the 35 stage × layer cells, no single benchmark does so in more than 10, and the agent post-training stage and the cluster layer have no task environment that runs the real system at all. The data stage has one, a data-selection task in AutoLab that is graded through a real fine-tune.
- The specialized benchmarks provide substantial training and inference coverage, but their task environments largely test individual components or bounded experiments, leaving gaps in coordinated, end-to-end workflows. SWE-Serve is the sharpest case: 21 of its 53 verifiers run a real SGLang engine, every one on a single GPU, and its tasks make up 53 of the 92 in the figure's serve column.
Where could new task environments come from?
Public training reports and run records describe the work involved in developing LLMs, while widely used open-source frameworks provide implementations across the lifecycle. Together, they offer a starting point for extending existing benchmarks with tasks grounded in real training and deployment workflows. The open side is further along than it may seem. Several open models now compete with frontier proprietary ones and were trained on open-source frameworks, GLM-5 on slime and Nemotron 3 on NeMo RL; Xiaomi even streamed MiMo-V2.6's RL run to a public dashboard. Reviewing these materials, the full cycle falls into the same five stages we audit against, from data through pretraining and post-training to agent post-training and deployment.
Table view
Where the public code lives
Code and open artifacts for each stage. Every stage has open code; what is thin is any account of operating that code across nodes and through failures. Each chip links to its source.
Code and artifacts, stage by stage
The materials for harder task environments therefore exist; what is missing is environments that run them.
What a good RSI benchmark needs
Section 02 showed that the task environments we audited cover little of end-to-end LLM training, and that the real system is never executed for agent post-training or cluster work, and only once for data. That is not carelessness on the benchmark builders' part. Training a production model takes large data, many GPUs across several nodes and long runs whose failures are distributed and slow to surface, and a task environment that must be built, run and graded in hours on one machine cannot hold most of that. The ceiling is the container, not the task author. A benchmark that measures end-to-end LLM training needs two things the current ones lack.
Tasks shaped like the work. The Terminal-Bench series (Terminal-Bench 2.1, 4.0 and Terminal-Bench-Science) shows that a community can build frontier-level task environments for terminal work, release after release. Harbor, the framework that grew out of that community, has become a common standard for building and running task environments. The same path is open for LLM training, and it has begun: MLS-Bench's 28 training and serving task environments, RSI-Exam's kernel and post-training environments and PostTrainBench's ten-H100-hour post-training task all appeared within a year, and MLS-Bench and RSI-Exam already ship in Harbor's format. What is still missing are task environments shaped like the work itself, where the agent must move a real training or serving system through a realistic problem, such as recovering a run after a failed node, keeping an asynchronous RL loop within a staleness budget, or publishing weights to a fleet of rollout engines, rather than editing one function behind a fixed harness.
Environments sized like the work. The environments in the figure are small by construction. Of its 133 task environments,2The 156 task environments that touch the lifecycle, with PostTrainBench's 28 variants and InferenceBench's 4 scenarios counted once each, plus seven adjacent-systems environments: five DeepSWE cluster-tooling tasks and the WeirdML and EdgeBench families, which reach related cells only. 44 declare no GPU at all, 60 declare one, 25 declare two to four, three declare five to eight, one (the EdgeBench family) does not document hardware per task, and none spans more than one node. Terminal-Bench 2.1 and DeepSWE declare zero GPUs for every task environment, Terminal-Bench 4.0 gives three of its 66 a single H100, and SWE-Serve caps every one of its 53 at one H100 by design. The largest environment anywhere in the set is RE-Bench's eight H100s, while the smallest documented pretraining run among the twelve pipelines in section 03, SmolLM3's, used 384 H100s.
So the environments are small, but not because a single node cannot be had. Harbor's task format declares GPU counts, and Modal and Daytona now give each trial its own sandbox with up to eight GPUs, metered while it runs and deleted afterwards; MLS-Bench already runs 8×H100 task environments on Daytona. What is still hard to ask for is several nodes. Harbor has no field for it, Daytona sandboxes stop at one machine, and Modal's RDMA clusters of up to 64 GPUs, still in beta, run as functions rather than as sandboxes an agent can be handed. On Daytona, a GPU task must also fit in one container, without the services a real training stack runs beside it. The next step is public environments that hold a small cluster for hours, expose real interconnects and schedulers, and grade what happens to a training run, not only what a unit test sees.
Early attempts. In the past two months, several projects have begun to test whether AI systems can do model development on their own. The proprietary Vals RSI Index scores models on four open-ended LLM research tasks, and its post-training task gives a single agent two eight-GPU nodes for 30 hours. RSI Arena, a live competition, gives eight agents the same 30B base model, 1,000 GPU-hours and 144 hours each to post-train it on 64 GPUs under Slurm, and judges the results by human evaluation and held-out tests. Two public efforts are still collecting task environments: the OpenRSI Index, a preview whose harness can spread a task over several nodes under Slurm or LSF, although no published run does, and Scale AI's RSI Bench, which plans to select its first 50 by 1 November 2026. We have not audited any of the four: the first two publish no task environments, and the last two are still changing.
Put together, a good RSI benchmark hands an agent a real training or serving system at realistic scale, breaks it the way production breaks, and grades the run that comes out: the quality of the model, the throughput of the job, whether it recovered. That is the capability RSI depends on, and every piece needed to test it, from the stacks to the harnesses, is now in the open.
Acknowledgements
Citation
Cited as:
Zhu, Jiacheng, Yichuan Wang and Zhifei Li. (Sep 2026). How to benchmark RSI. Artificial Evolution Lab, Vol. 1. https://artificialevolution.org/vol/1/
Or
@article{aelab2026benchmarkrsi,
title = {How to benchmark RSI},
author = {Zhu, Jiacheng and Wang, Yichuan and Li, Zhifei},
journal = {Artificial Evolution Lab},
volume = {1},
year = {2026},
month = {Sep},
url = {https://artificialevolution.org/vol/1/}
}
References
Sources are linked where they are mentioned and collected here in order of first appearance. For benchmarks, the commit at which we read their task environments is given in grey.
- Anthropic. “Measurements for understanding the pace of AI development inside frontier labs.” Anthropic (Sep 2026).
- OpenAI. “Research acceleration: The view inside OpenAI.” OpenAI (Sep 2026).
- Z.ai. “Toward Recursive Self-Improvement: How GLM Built Its Own Inference Infrastructure.” Z.ai blog (Sep 2026).
- Xiang Deng, et al. “SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?” arXiv:2509.16941 (2025).
- Terminal-Bench team. “Terminal-Bench 4.0.” tbench.ai (2026). tasks read at harbor-framework/terminal-bench@452bf30
- Terminal-Bench team. “Terminal-Bench 2.1.” tbench.ai (2026). tasks read at harbor-framework/terminal-bench-2-1@7131e43
- Datacurve. “DeepSWE.” deepswe.datacurve.ai (2026). tasks read at datacurve-ai/deep-swe@0b9fabb
- John Yang, et al. “ProgramBench: Can Language Models Rebuild Programs From Scratch?” arXiv:2605.03546 (2026).
- Harbor. “Harbor documentation: datasets, sandboxes and the task format.” docs.harborframework.com (2026).
- OpenAI. “OpenAI and Hugging Face partner to address security incident during model evaluation.” OpenAI (Jul 2026).
- An Yang, et al. “Qwen3 Technical Report.” arXiv:2505.09388 (2025).
- Kimi Team. “Kimi K2: Open Agentic Intelligence.” arXiv:2507.20534 (2025).
- GLM-5 Team. “GLM-5: from Vibe Coding to Agentic Engineering.” arXiv:2602.15763 (2026).
- Ai2. “Olmo 3: Charting a path through the model flow to lead open-source AI.” Ai2 blog (Nov 2025).
- Hugging Face. “SmolLM3: smol, multilingual, long-context reasoner.” Hugging Face blog (Jul 2025).
- MLS-Bench team. “MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI.” mls-bench.com (2026). tasks read at Imbernoulli/MLS-Bench@4b1fd67
- Hjalmar Wijk, et al. “RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts.” arXiv:2411.15114 (2024). tasks read at METR/RE-Bench@93b9806
- RSI-Exam team. “RSI-Exam: Benchmarking Recursive Self-Improvement through Executable Research.” rsi-exam.ai (2026). tasks read at huggingface.co/datasets/RSI-Exam/RSI-Exam@23ab61b
- Zhangchen Xu, et al. “AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?” arXiv:2606.05080 (2026). tasks read at autolabhq/autolab@4127da3
- NVIDIA. “SWE-Serve: an agentic benchmark of 53 production inference-engineering tasks.” GitHub (2026). tasks read at NVIDIA/swe-serve@e4286ec
- Jehyeok Yeon, et al. “InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents.” arXiv:2607.20468 (2026).
- Ben Rank, et al. “PostTrainBench: Can LLM Agents Automate LLM Post-Training?” arXiv:2603.08640 (2026).
- Deyao Zhu, et al. “EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments.” arXiv:2607.05155 (2026).
- Håvard Tveit Ihle. “WeirdML v3.” htihle.github.io (2026).
- THUDM. “slime: an LLM post-training framework for RL scaling.” GitHub (2025).
- NVIDIA. “NeMo RL: a scalable toolkit for efficient model reinforcement learning.” GitHub (2025).
- Xiaomi LLM-Core. “MiMo-V2.6 RL training dashboard.” mimo.xiaomi.com (Sep 2026).
- GLM-4.5 Team. “GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.” arXiv:2508.06471 (2025).
- DeepSeek-AI. “DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence.” arXiv:2606.19348 (2026).
- Aili Chen, et al. “The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence.” arXiv:2605.26494 (2026).
- Xiaomi LLM-Core. “MiMo-V2.6 Technical Report.” Hugging Face (2026).
- NVIDIA. “Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning.” arXiv:2604.12374 (2026).
- Open Athena. “Marin 535B-A23B launch note.” openathena.ai (Sep 2026).
- Mercor. “Training frontier knowledge work agents: A 397B RL training guide with SkyRL.” Mercor blog (Sep 2026).
- NovaSky. “SkyRL: A modular full-stack RL library for LLMs.” GitHub (2025).
- Terminal-Bench team. “Terminal-Bench-Science: Evaluating AI agents on research workflows across scientific domains.” GitHub (2026).
- Modal. “Sandbox resources, GPU acceleration and multi-node clusters.” modal.com docs (2026).
- Daytona. “Sandboxes: GPU sandboxes.” daytona.io docs (2026).
- MLS-Bench team. “Daytona: make GPU sandboxes ephemeral so organizations that require it accept them.” Pull request #98, GitHub (Sep 2026).
- Oliver Chen, et al. “Vals RSI Index.” Vals AI (2026).
- Bake AI. “RSI Arena.” rsiarena.org (Sep 2026).
- OpenRSI Foundation. “OpenRSI Index, Preview v0.1.” index.openrsi.foundation (Sep 2026).
- Scale AI. “RSI Bench.” labs.scale.com (2026).