MiMo-V2.6 release review (Xiaomi LLM-Core, 21–22 September 2026)

Reviewed 2026-09-21/22. Every fact below was read directly from a primary source in this session (tag spot-checked): the 44-page technical report (PDF created 2026-09-21, Other_company_recipe/Xiaomi_Tech_Report/MiMo_V2_6_technical_report.pdf, also on Hugging Face), the blog article (mimo.xiaomi.com/mimo-v2-6, served as an iframe article.html, dated "September 22nd, 2026"), the news post on mimo.mi.com ("Update Time September 21, 2026"), the three Hugging Face model cards, the live RL dashboard (mimo.xiaomi.com/rl/mimo-v26; overview, metrics tag list, batch composition, notices), the XiaomiMiMo GitHub organization (API listing and the three mimo-oss commits), the MiMo-V2-Flash report and README, and the abstracts of the building-block papers (R3, MOPD, Muown, DFlash, CodeMidas). Press coverage (VentureBeat, Forkast, TestingCatalog, a Substack post) is tagged secondary and used only where it adds something the primary sources do not say. Extracted source text is kept under sources/mimo_v26/.

1. What was released

Artifact Status on 2026-09-21 (evening UTC)
MiMo-V2.6-Pro-RL weights Hugging Face + ModelScope, MIT. 1.02T total / 42B active sparse MoE; 70 layers (60 SWA + 10 GA), hidden 6144, 384 experts / 8 active, no shared experts; 1M context; omnimodal (MiMo-ViT 681M; audio tokenizer 308M + patch encoder 127M); 5-layer SWA DFlash-style MTP drafter predicting 7 tokens.
MiMo-V2.6-Flash-RL weights Same license and channels. 310B total (report Table 1) or 309B (model card, news) / 15B active; 48 layers (39 + 9), hidden 4096, 256 experts / 8 active.
MiMo-V2.6-Pro-UltraSpeed API and desktop mode, "up to 20× faster output" at 10× the Pro price; no weights.
MiMo-V2.6-Distill-Qwen-9B Hugging Face. SFT of Qwen3.5-9B on 77.4B tokens of MiMo-generated data (27.2B loss tokens: code 23.2B, cyber 11.0B, general 22.0B, visual 21.2B). Released "as a starting point for open research in agentic reinforcement learning".
Technical report 44 pages, in the model repos. Cites its own building blocks: R3 (arXiv 2510.11370), MOPD (2606.30406), Muown (2605.10797), DFlash (2602.06036), CodeMidas (2609.22068), Humming MXFP4 kernels (vllm-project/humming v0.1.15), AsyncFlow/TransferQueue (2507.01663), Kimi k1.5 partial rollout, DAPO dynamic sampling, MAI-Thinking-1 top-p replay, DeepSeek-V3.2.
"7k+ high-quality RL task environments" Announced (news post, report Table 5: ~3k software engineering with executable tests, ~1k vulnerability reproduction with rule checks, ~1k knowledge work with rubric judging, ~2k web development with visual grading, plus ~1k music tasks). Not yet published: the HF collection XiaomiMiMo/mimo-v26 holds only the three models, the org lists no new datasets, and the report's only footnote link is the Distill model.
"End-to-end RL training framework" Published as three forks on GitHub, each with a mimo-oss default branch created 2026-09-21: XiaomiMiMo/verl (on verl v0.9.0.dev commit 2781c1e; one commit, +30,672 lines), XiaomiMiMo/uni-agent (on verl-project/uni-agent eac7985; +5,002 lines), XiaomiMiMo/mimoagent (fork of mini-swe-agent v1.9.0, "largely rewritten"). The news post says the framework is "built on verl, uni-agent and mini-swe-agent".
"Composable mini-harnesses" The MiMoAgent agents (default, bashonly-agent, cc-agent with the Claude Code tool catalogue, codex-agent with the Codex catalogue, mimocode-agent) plus black-box CLI adapters; the verl branch adds "mimoagent SWE, mini-swe-agent, the Claude Code and Codex Responses harnesses, and the uni-agent blackbox path".
API pricing (per million tokens, cache miss / output) Flash $0.14 / $0.28; Pro $0.435 / $0.87; UltraSpeed $4.35 / $8.7; cache-hit input $0.0028 / $0.0036 / $0.036. Unchanged from V2.5.

The production RL system described in report §6 (SGLang + Megatron-LM with Xiaomi's Harness Pool, Payload Porter, Sample Mixer) is not what was open-sourced. The open suite is a verl-based reconstruction of the recipes and harness mechanics. Keep the two apart when citing.

2. The lifecycle, in the audit's six stages

3. The RL run, as documented

Setting Value
Algorithm GRPO with prompt-mean loss aggregation; token-level importance ratio against the rollout engine's log-probabilities; decoupled clip bounds for positive and negative advantages initialized to [0.2, 5.0] and retuned at runtime from policy entropy
Batch 1,568 prompts × 16 rollouts = 25K sequences per step; 2.7–3.7B tokens per step (about 110–150K tokens per sequence); context up to 1M
Asynchrony Fully asynchronous with partial rollouts at a staleness of 4; partial rollouts keep their original inference log-probabilities and re-prefill after each policy update
Optimizer Muown, lr 3e-6, no weight decay or warmup, grad clip 1.0; Muon momentum 0.95 (Nesterov), 10 Newton–Schulz iterations, update scale 0.5; Adam β1 = β2 = 0.95; FP32 master weights and row state carried over from SFT; MoE router frozen
Task mix (report) agentic + competitive coding 68%, general tool use 12%, aesthetic design 13%, context following 3%, cybersecurity 4%
Steps and cost 30 steps each. Pro: $2.62M, 123.1 h, 75B tokens, 753k samples. Flash: $0.854M, 81.8 h, 81.4B tokens, 753k samples. Total dashboard cost $3,474,715. Pro cost split: rollout 43.8%, training 43.5%, grader 12.7%; Flash 44.9 / 40.9 / 14.2
Dashboard start/stop Pro started 2026-09-15 10:32 UTC, stopped 2026-09-20 (5 d 07:29); Flash started 15:16 UTC, stopped 2026-09-19 (3 d 11:05); streaming since 2026-09-16 04:00 UTC
Progress DeepSWE v1.1 avg@3 (mini-swe-agent harness): Pro 58.4 → 72.57, Flash 48.7 → 65.68; training-task pass rate (dynsam/avg@n) Pro 0.633 (+0.068 vs step 1), Flash 0.644 (+0.130)
Grading Groupwise Reward Synthesis (offline rubrics; R = R_test · S_sol · S_beh) on high-pass-rate code tasks; Groupwise Advantage Redistribution (online SFT-trained agentic grader ranks passing patches on five dimensions, zeroes confirmed hacks, redistributes positive advantage with a capped common factor); group-relative length penalty; segment-level penalties for format and tool-call errors with conserved advantage mass
Reward-hacking defense Mid-training self-correction data; environment preparation (artifact, cache and git cleanup, container-level network isolation); a hack agent iterated until no exploit remains; offline trajectory audits during training; confirmed-hack share below 2% throughout
Harnesses Four training mini-harnesses per domain; held-out harness (codex, claude code, mini-swe-agent) mean pass@1 on DeepSWE rose from about 50% to 66%; dashboard step 30 shows 23 trained harnesses over 22,729 rollouts, the largest at 18.9%

4. Infrastructure, component by component (report §6 + dashboard)

5. Operational evidence (the part RL-Infra-Bench should mine)

Report Figure 12 and the dashboard notices agree and complement each other:

Incident Source Detail
GPU memory double-bit errors report §5.5; dashboard "vram issue on one node" primary infrastructure failure class for both runs
Kubernetes failure in the cyber-task cluster report; dashboard "infra error on one of datasets was not correctly detected over the past ~3 hours" Flash restarted from step 15
Grader unreachable over the network report; dashboard Pro restarted after step 14; Xiaomi also "removed the cyber dataset from the upcoming pro run, since we observed some bad patterns in the rollout logs"
Inference failures report partial-rollout startup bias exhausted both the GPU and pinned host KV pools; one harness produced rollouts less than half as long as the others on the same dataset, skewing dispatch estimates for two steps
Training failure: GPU OOM from MoE imbalance report; dashboard "pro run restarted at step 17 due to a GPU OOM issue caused by expert load imbalance" an EP rank received over 30× the mean token load inside one micro-batch despite a balanced full batch; parallelism adjusted to reduce activation memory
Driver failure: CPU OOM during packing report late in Flash as longer sequences raised per-node data volume beyond host memory, even with packing distributed across nodes
Router load collapse report Figure 11 with a trainable router, over 20 steps CV 0.78 → 2.0, peak load 6× → 16×, cold experts 0.5% → 22%; restoring router weights alone restored balance with unchanged benchmarks; frozen router keeps CV ≈ 0.7, peak ≈ 5.5×, cold ≈ 1%
Mid-run curriculum change dashboard "we filtered out tasks that are relatively easy for the current pro model"

Dashboard batch composition at Pro step 30: 25 sources; code 1,033 prompts (65.8%), visual 305 (19.5%), general 176 (11.3%), chat 54 (3.4%), cyber 0. Dynamic-sampler feed at Pro step 30: accepted 2,208 / 1,568 target, judged 2,155, pass 0.606 (n = 3,644), remaining 332 + 537 partial + 191 rewarding, prewarm 256. Flash step 31 (at stop): accepted 2,704 / 1,568, pass 0.547, with 1,001 partial. The streamed metric tags (2,077 in total) include partial/avg_staleness, dynsam/infra_error/seq_rate, env/active, timing_s/outer_gen and timing_s/trainer_ops; the charts render on canvas and were not read numerically here.

6. Benchmarks (Table 3, baselines at maximum reasoning effort)

Benchmark Pro Flash V2.5-Pro Opus 5 GPT-5.6 Sol Fable 5
DeepSWE v1.1 71.9 67.9 19.0 74.0 73.0 70.0
ProgramBench 26.5 26.0 12.5 37.0 25.0 33.0
Terminal-Bench 4.0 34.9 28.8 1.5 49.0 39.9 42.4
Terminal-Bench 2.1 89.9 87.6 65.2 89.1 88.8 84.3
AutomationBench v1.0.6 53.1 52.3 16.0 50.3 45.8 46.2
Toolathlon-Verified 76.9 73.6 49.1 80.6 74.9 77.9
OSWorld-Verified 82.0 80.8 — 83.4 83.0 86.0
CyberGym (corrected environments) 94.0 95.1 40.0 — — —
ExploitBench 47.9 25.3 16.6 70.0 78.5 78.0

Two readings matter for the audit. The 55-point gap between Terminal-Bench 2.1 (saturated, 89.9) and 4.0 (34.9) confirms that the two versions must never share a score series. And the MiMo Code Bench, MiMo Cyber Bench and MiMo Visual Coding numbers are internal and unreproducible. The Distill-9B experiments (Table 6) show RL on the released environments lifting SWE-bench Verified 61.1 → 66.2 and Terminal-Bench 2.1 37.1 → 52.8, and multi-harness RL improving all 21 dataset–harness pairs.

7. Discrepancies and open questions

  1. Flash size: 310B (report) versus 309B (model card, news). Table 1 prints "42T" active parameters for Pro; the text and card say 42B.
  2. Tokens per step: report 2.7–3.7B; blog and news 3.5–3.7B; dashboard step 30 shows 3.43B (Pro) and 3.7B (Flash).
  3. Cyber share: the report says 4% of RL tasks; the dashboard shows cyber removed from Pro after the grader incident (0.0% at step 30) while Flash step 31 still accepted 145 cyber groups against a target of 64.
  4. Cluster size: "thousands of GPUs" and dollar costs only; no GPU type or count. The Forkast estimate of about $432k per day is derived from the dashboard, not stated by Xiaomi (secondary).
  5. Environments: announced as open, not yet published; framework code is. The report's abstract says "we open-source the training dynamics, RL environments, and RL framework"; treat the environments as pending until a dataset appears.
  6. Forkast's claim of a "Claude Distill Requests: hidden" dashboard field was not present in the dashboard text read here (secondary, unverified).
  7. The report attributes the CyberGym score to "corrected" evaluation environments (footnote 1); the correction method is the report's own sanitizer-report matching, so the number is not comparable to the public CyberGym leaderboard.

8. Relevance to the coverage audit and RL-Infra-Bench

MiMo-V2.6 becomes the eighteenth documented pipeline on the coverage page and the most detailed first-party account so far of mixed-task asynchronous agentic RL operations: an honest failure ledger with a timeline, a cost decomposition across rollout, training and grader, per-source scheduling formulas, and a streamed metric catalogue. It grounds every RL-Infra-Bench mechanism family with a human incident:

Candidate source-backed task shapes (each scales down by mechanism, not size): "the grader endpoint becomes unreachable mid-step; the run must stall or fail cleanly without training on unjudged groups"; "one EP rank receives 30× the mean tokens in a micro-batch; diagnose from memory telemetry and fix by repacking or parallelism"; "after a restart the dispatcher's length priors are biased by short rollouts and the KV pools are exhausted"; "router drift produces cold experts; detect via CV, peak load and cold-expert fraction and freeze the router"; "one harness yields rollouts half as long as its peers; the sample mixer must keep per-source shares". The open verl mimo-oss branch, uni-agent gateway and MiMoAgent (with Kubernetes and Modal backends and DeepSWE, TerminalBench and ProgramBench adapters) are also a plausible pipeline pack for RL-Infra-Bench once the environments ship.