Reviewed 2026-09-21/22. Every fact below was read directly from a primary source in this session (tag spot-checked): the 44-page technical report (PDF created 2026-09-21, Other_company_recipe/Xiaomi_Tech_Report/MiMo_V2_6_technical_report.pdf, also on Hugging Face), the blog article (mimo.xiaomi.com/mimo-v2-6, served as an iframe article.html, dated "September 22nd, 2026"), the news post on mimo.mi.com ("Update Time September 21, 2026"), the three Hugging Face model cards, the live RL dashboard (mimo.xiaomi.com/rl/mimo-v26; overview, metrics tag list, batch composition, notices), the XiaomiMiMo GitHub organization (API listing and the three mimo-oss commits), the MiMo-V2-Flash report and README, and the abstracts of the building-block papers (R3, MOPD, Muown, DFlash, CodeMidas). Press coverage (VentureBeat, Forkast, TestingCatalog, a Substack post) is tagged secondary and used only where it adds something the primary sources do not say. Extracted source text is kept under sources/mimo_v26/.
| Artifact | Status on 2026-09-21 (evening UTC) |
|---|---|
| MiMo-V2.6-Pro-RL weights | Hugging Face + ModelScope, MIT. 1.02T total / 42B active sparse MoE; 70 layers (60 SWA + 10 GA), hidden 6144, 384 experts / 8 active, no shared experts; 1M context; omnimodal (MiMo-ViT 681M; audio tokenizer 308M + patch encoder 127M); 5-layer SWA DFlash-style MTP drafter predicting 7 tokens. |
| MiMo-V2.6-Flash-RL weights | Same license and channels. 310B total (report Table 1) or 309B (model card, news) / 15B active; 48 layers (39 + 9), hidden 4096, 256 experts / 8 active. |
| MiMo-V2.6-Pro-UltraSpeed | API and desktop mode, "up to 20× faster output" at 10× the Pro price; no weights. |
| MiMo-V2.6-Distill-Qwen-9B | Hugging Face. SFT of Qwen3.5-9B on 77.4B tokens of MiMo-generated data (27.2B loss tokens: code 23.2B, cyber 11.0B, general 22.0B, visual 21.2B). Released "as a starting point for open research in agentic reinforcement learning". |
| Technical report | 44 pages, in the model repos. Cites its own building blocks: R3 (arXiv 2510.11370), MOPD (2606.30406), Muown (2605.10797), DFlash (2602.06036), CodeMidas (2609.22068), Humming MXFP4 kernels (vllm-project/humming v0.1.15), AsyncFlow/TransferQueue (2507.01663), Kimi k1.5 partial rollout, DAPO dynamic sampling, MAI-Thinking-1 top-p replay, DeepSeek-V3.2. |
| "7k+ high-quality RL task environments" | Announced (news post, report Table 5: ~3k software engineering with executable tests, ~1k vulnerability reproduction with rule checks, ~1k knowledge work with rubric judging, ~2k web development with visual grading, plus ~1k music tasks). Not yet published: the HF collection XiaomiMiMo/mimo-v26 holds only the three models, the org lists no new datasets, and the report's only footnote link is the Distill model. |
| "End-to-end RL training framework" | Published as three forks on GitHub, each with a mimo-oss default branch created 2026-09-21: XiaomiMiMo/verl (on verl v0.9.0.dev commit 2781c1e; one commit, +30,672 lines), XiaomiMiMo/uni-agent (on verl-project/uni-agent eac7985; +5,002 lines), XiaomiMiMo/mimoagent (fork of mini-swe-agent v1.9.0, "largely rewritten"). The news post says the framework is "built on verl, uni-agent and mini-swe-agent". |
| "Composable mini-harnesses" | The MiMoAgent agents (default, bashonly-agent, cc-agent with the Claude Code tool catalogue, codex-agent with the Codex catalogue, mimocode-agent) plus black-box CLI adapters; the verl branch adds "mimoagent SWE, mini-swe-agent, the Claude Code and Codex Responses harnesses, and the uni-agent blackbox path". |
| API pricing (per million tokens, cache miss / output) | Flash $0.14 / $0.28; Pro $0.435 / $0.87; UltraSpeed $4.35 / $8.7; cache-hit input $0.0028 / $0.0036 / $0.036. Unchanged from V2.5. |
The production RL system described in report §6 (SGLang + Megatron-LM with Xiaomi's Harness Pool, Payload Porter, Sample Mixer) is not what was open-sourced. The open suite is a verl-based reconstruction of the recipes and harness mechanics. Keep the two apart when citing.
--reasoning-parser mimo; Flash TP 8, DP 2; vLLM at TP 8 (Pro) and TP 4 (Flash) with a mimov25-cu129 image. Rollouts and, by the model card's Humming reference, serving use MXFP4 experts. No production serving numbers are published; a separate news item reports a 1T model exceeding 1,000 tokens/s with TileRT, which is presumably the UltraSpeed path, but the release does not say so.| Setting | Value |
|---|---|
| Algorithm | GRPO with prompt-mean loss aggregation; token-level importance ratio against the rollout engine's log-probabilities; decoupled clip bounds for positive and negative advantages initialized to [0.2, 5.0] and retuned at runtime from policy entropy |
| Batch | 1,568 prompts × 16 rollouts = 25K sequences per step; 2.7–3.7B tokens per step (about 110–150K tokens per sequence); context up to 1M |
| Asynchrony | Fully asynchronous with partial rollouts at a staleness of 4; partial rollouts keep their original inference log-probabilities and re-prefill after each policy update |
| Optimizer | Muown, lr 3e-6, no weight decay or warmup, grad clip 1.0; Muon momentum 0.95 (Nesterov), 10 Newton–Schulz iterations, update scale 0.5; Adam β1 = β2 = 0.95; FP32 master weights and row state carried over from SFT; MoE router frozen |
| Task mix (report) | agentic + competitive coding 68%, general tool use 12%, aesthetic design 13%, context following 3%, cybersecurity 4% |
| Steps and cost | 30 steps each. Pro: $2.62M, 123.1 h, 75B tokens, 753k samples. Flash: $0.854M, 81.8 h, 81.4B tokens, 753k samples. Total dashboard cost $3,474,715. Pro cost split: rollout 43.8%, training 43.5%, grader 12.7%; Flash 44.9 / 40.9 / 14.2 |
| Dashboard start/stop | Pro started 2026-09-15 10:32 UTC, stopped 2026-09-20 (5 d 07:29); Flash started 15:16 UTC, stopped 2026-09-19 (3 d 11:05); streaming since 2026-09-16 04:00 UTC |
| Progress | DeepSWE v1.1 avg@3 (mini-swe-agent harness): Pro 58.4 → 72.57, Flash 48.7 → 65.68; training-task pass rate (dynsam/avg@n) Pro 0.633 (+0.068 vs step 1), Flash 0.644 (+0.130) |
| Grading | Groupwise Reward Synthesis (offline rubrics; R = R_test · S_sol · S_beh) on high-pass-rate code tasks; Groupwise Advantage Redistribution (online SFT-trained agentic grader ranks passing patches on five dimensions, zeroes confirmed hacks, redistributes positive advantage with a capped common factor); group-relative length penalty; segment-level penalties for format and tool-call errors with conserved advantage mass |
| Reward-hacking defense | Mid-training self-correction data; environment preparation (artifact, cache and git cleanup, container-level network isolation); a hack agent iterated until no exploit remains; offline trajectory audits during training; confirmed-hack share below 2% throughout |
| Harnesses | Four training mini-harnesses per domain; held-out harness (codex, claude code, mini-swe-agent) mean pass@1 on DeepSWE rose from about 50% to 66%; dashboard step 30 shows 23 trained harnesses over 22,729 rollouts, the largest at 18.9% |
train_infer_diff/new_infer/kl metric is streamed on the dashboard.Report Figure 12 and the dashboard notices agree and complement each other:
| Incident | Source | Detail |
|---|---|---|
| GPU memory double-bit errors | report §5.5; dashboard "vram issue on one node" | primary infrastructure failure class for both runs |
| Kubernetes failure in the cyber-task cluster | report; dashboard "infra error on one of datasets was not correctly detected over the past ~3 hours" | Flash restarted from step 15 |
| Grader unreachable over the network | report; dashboard | Pro restarted after step 14; Xiaomi also "removed the cyber dataset from the upcoming pro run, since we observed some bad patterns in the rollout logs" |
| Inference failures | report | partial-rollout startup bias exhausted both the GPU and pinned host KV pools; one harness produced rollouts less than half as long as the others on the same dataset, skewing dispatch estimates for two steps |
| Training failure: GPU OOM from MoE imbalance | report; dashboard "pro run restarted at step 17 due to a GPU OOM issue caused by expert load imbalance" | an EP rank received over 30× the mean token load inside one micro-batch despite a balanced full batch; parallelism adjusted to reduce activation memory |
| Driver failure: CPU OOM during packing | report | late in Flash as longer sequences raised per-node data volume beyond host memory, even with packing distributed across nodes |
| Router load collapse | report Figure 11 | with a trainable router, over 20 steps CV 0.78 → 2.0, peak load 6× → 16×, cold experts 0.5% → 22%; restoring router weights alone restored balance with unchanged benchmarks; frozen router keeps CV ≈ 0.7, peak ≈ 5.5×, cold ≈ 1% |
| Mid-run curriculum change | dashboard | "we filtered out tasks that are relatively easy for the current pro model" |
Dashboard batch composition at Pro step 30: 25 sources; code 1,033 prompts (65.8%), visual 305 (19.5%), general 176 (11.3%), chat 54 (3.4%), cyber 0. Dynamic-sampler feed at Pro step 30: accepted 2,208 / 1,568 target, judged 2,155, pass 0.606 (n = 3,644), remaining 332 + 537 partial + 191 rewarding, prewarm 256. Flash step 31 (at stop): accepted 2,704 / 1,568, pass 0.547, with 1,001 partial. The streamed metric tags (2,077 in total) include partial/avg_staleness, dynsam/infra_error/seq_rate, env/active, timing_s/outer_gen and timing_s/trainer_ops; the charts render on canvas and were not read numerically here.
| Benchmark | Pro | Flash | V2.5-Pro | Opus 5 | GPT-5.6 Sol | Fable 5 |
|---|---|---|---|---|---|---|
| DeepSWE v1.1 | 71.9 | 67.9 | 19.0 | 74.0 | 73.0 | 70.0 |
| ProgramBench | 26.5 | 26.0 | 12.5 | 37.0 | 25.0 | 33.0 |
| Terminal-Bench 4.0 | 34.9 | 28.8 | 1.5 | 49.0 | 39.9 | 42.4 |
| Terminal-Bench 2.1 | 89.9 | 87.6 | 65.2 | 89.1 | 88.8 | 84.3 |
| AutomationBench v1.0.6 | 53.1 | 52.3 | 16.0 | 50.3 | 45.8 | 46.2 |
| Toolathlon-Verified | 76.9 | 73.6 | 49.1 | 80.6 | 74.9 | 77.9 |
| OSWorld-Verified | 82.0 | 80.8 | — | 83.4 | 83.0 | 86.0 |
| CyberGym (corrected environments) | 94.0 | 95.1 | 40.0 | — | — | — |
| ExploitBench | 47.9 | 25.3 | 16.6 | 70.0 | 78.5 | 78.0 |
Two readings matter for the audit. The 55-point gap between Terminal-Bench 2.1 (saturated, 89.9) and 4.0 (34.9) confirms that the two versions must never share a score series. And the MiMo Code Bench, MiMo Cyber Bench and MiMo Visual Coding numbers are internal and unreproducible. The Distill-9B experiments (Table 6) show RL on the released environments lifting SWE-bench Verified 61.1 → 66.2 and Terminal-Bench 2.1 37.1 → 52.8, and multi-harness RL improving all 21 dataset–harness pairs.
MiMo-V2.6 becomes the eighteenth documented pipeline on the coverage page and the most detailed first-party account so far of mixed-task asynchronous agentic RL operations: an honest failure ledger with a timeline, a cost decomposition across rollout, training and grader, per-source scheduling formulas, and a streamed metric catalogue. It grounds every RL-Infra-Bench mechanism family with a human incident:
Candidate source-backed task shapes (each scales down by mechanism, not size): "the grader endpoint becomes unreachable mid-step; the run must stall or fail cleanly without training on unjudged groups"; "one EP rank receives 30× the mean tokens in a micro-batch; diagnose from memory telemetry and fix by repacking or parallelism"; "after a restart the dispatcher's length priors are biased by short rollouts and the KV pools are exhausted"; "router drift produces cold experts; detect via CV, peak load and cold-expert fraction and freeze the router"; "one harness yields rollouts half as long as its peers; the sample mixer must keep per-source shares". The open verl mimo-oss branch, uni-agent gateway and MiMoAgent (with Kubernetes and Modal backends and DeepSWE, TerminalBench and ProgramBench adapters) are also a plausible pipeline pack for RL-Infra-Bench once the environments ship.