# MiMo-V2.6 release review (Xiaomi LLM-Core, 21–22 September 2026) Reviewed 2026-09-21/22. Every fact below was read directly from a primary source in this session (tag **spot-checked**): the 44-page technical report (PDF created 2026-09-21, `Other_company_recipe/Xiaomi_Tech_Report/MiMo_V2_6_technical_report.pdf`, also on Hugging Face), the blog article (mimo.xiaomi.com/mimo-v2-6, served as an iframe `article.html`, dated "September 22nd, 2026"), the news post on mimo.mi.com ("Update Time September 21, 2026"), the three Hugging Face model cards, the live RL dashboard (mimo.xiaomi.com/rl/mimo-v26; overview, metrics tag list, batch composition, notices), the XiaomiMiMo GitHub organization (API listing and the three `mimo-oss` commits), the MiMo-V2-Flash report and README, and the abstracts of the building-block papers (R3, MOPD, Muown, DFlash, CodeMidas). Press coverage (VentureBeat, Forkast, TestingCatalog, a Substack post) is tagged **secondary** and used only where it adds something the primary sources do not say. Extracted source text is kept under `sources/mimo_v26/`. ## 1. What was released | Artifact | Status on 2026-09-21 (evening UTC) | |---|---| | MiMo-V2.6-Pro-RL weights | Hugging Face + ModelScope, MIT. 1.02T total / 42B active sparse MoE; 70 layers (60 SWA + 10 GA), hidden 6144, 384 experts / 8 active, no shared experts; 1M context; omnimodal (MiMo-ViT 681M; audio tokenizer 308M + patch encoder 127M); 5-layer SWA DFlash-style MTP drafter predicting 7 tokens. | | MiMo-V2.6-Flash-RL weights | Same license and channels. 310B total (report Table 1) or 309B (model card, news) / 15B active; 48 layers (39 + 9), hidden 4096, 256 experts / 8 active. | | MiMo-V2.6-Pro-UltraSpeed | API and desktop mode, "up to 20× faster output" at 10× the Pro price; no weights. | | MiMo-V2.6-Distill-Qwen-9B | Hugging Face. SFT of Qwen3.5-9B on 77.4B tokens of MiMo-generated data (27.2B loss tokens: code 23.2B, cyber 11.0B, general 22.0B, visual 21.2B). Released "as a starting point for open research in agentic reinforcement learning". | | Technical report | 44 pages, in the model repos. Cites its own building blocks: R3 (arXiv 2510.11370), MOPD (2606.30406), Muown (2605.10797), DFlash (2602.06036), CodeMidas (2609.22068), Humming MXFP4 kernels (vllm-project/humming v0.1.15), AsyncFlow/TransferQueue (2507.01663), Kimi k1.5 partial rollout, DAPO dynamic sampling, MAI-Thinking-1 top-p replay, DeepSeek-V3.2. | | "7k+ high-quality RL task environments" | Announced (news post, report Table 5: ~3k software engineering with executable tests, ~1k vulnerability reproduction with rule checks, ~1k knowledge work with rubric judging, ~2k web development with visual grading, plus ~1k music tasks). **Not yet published**: the HF collection `XiaomiMiMo/mimo-v26` holds only the three models, the org lists no new datasets, and the report's only footnote link is the Distill model. | | "End-to-end RL training framework" | Published as three forks on GitHub, each with a `mimo-oss` default branch created 2026-09-21: `XiaomiMiMo/verl` (on verl v0.9.0.dev commit 2781c1e; one commit, +30,672 lines), `XiaomiMiMo/uni-agent` (on verl-project/uni-agent eac7985; +5,002 lines), `XiaomiMiMo/mimoagent` (fork of mini-swe-agent v1.9.0, "largely rewritten"). The news post says the framework is "built on verl, uni-agent and mini-swe-agent". | | "Composable mini-harnesses" | The MiMoAgent agents (`default`, `bashonly-agent`, `cc-agent` with the Claude Code tool catalogue, `codex-agent` with the Codex catalogue, `mimocode-agent`) plus black-box CLI adapters; the verl branch adds "mimoagent SWE, mini-swe-agent, the Claude Code and Codex Responses harnesses, and the uni-agent blackbox path". | | API pricing (per million tokens, cache miss / output) | Flash $0.14 / $0.28; Pro $0.435 / $0.87; UltraSpeed $4.35 / $8.7; cache-hit input $0.0028 / $0.0036 / $0.036. Unchanged from V2.5. | The production RL system described in report §6 (SGLang + Megatron-LM with Xiaomi's Harness Pool, Payload Porter, Sample Mixer) is **not** what was open-sourced. The open suite is a verl-based reconstruction of the recipes and harness mechanics. Keep the two apart when citing. ## 2. The lifecycle, in the audit's six stages - **Data.** Text (web, books, papers, code, STEM), vision (captioning, grounding, OCR, GUI, video, visual-coding) and audio (speech-text interleaving, ASR, captioning) corpora described; nothing released. The ViT was pretrained on more than 4T image tokens paired with a small LLM; the audio tokenizer on 20 million hours. - **Pretrain.** Two stages: text-only backbone, then joint omni training with the in-house encoders. Flash 48T tokens (26T text + 22T omni); Pro 30T (27T + 3T). Context 32K, extended to 256K mid-run. AdamW. - **Midtrain.** Explicitly "agent-centric mid-training": realistic agent trajectories across coding, general, visual and research tasks mixed with text, repository-level code and image/video/audio; 256K context for most of the compute, then 1M in a final stage. The optimizer is switched from AdamW to Muown (Muon with row-norm control) for hidden matrices, with no loss spike; MXFP4 quantization-aware training; reward-hacking "self-correction" alignment examples are added here. - **Post-train.** A short SFT, then one mixed RL run ("You Only RL Once"), then MOPD2 (multi-prefix multi-teacher on-policy distillation) to fold in hard-to-verify domains from mixRL and SFT teachers. - **Agentic RL** is the substance of the report (§4, §5, §6); see §3 and §4 below. - **Serve.** SGLang recommended: Pro on two nodes with TP 16, DP 2 with DP attention, EP 16 over DeepEP, multi-layer EAGLE-style speculative decoding, `--reasoning-parser mimo`; Flash TP 8, DP 2; vLLM at TP 8 (Pro) and TP 4 (Flash) with a `mimov25-cu129` image. Rollouts and, by the model card's Humming reference, serving use MXFP4 experts. No production serving numbers are published; a separate news item reports a 1T model exceeding 1,000 tokens/s with TileRT, which is presumably the UltraSpeed path, but the release does not say so. ## 3. The RL run, as documented | Setting | Value | |---|---| | Algorithm | GRPO with prompt-mean loss aggregation; token-level importance ratio against the rollout engine's log-probabilities; decoupled clip bounds for positive and negative advantages initialized to [0.2, 5.0] and retuned at runtime from policy entropy | | Batch | 1,568 prompts × 16 rollouts = 25K sequences per step; 2.7–3.7B tokens per step (about 110–150K tokens per sequence); context up to 1M | | Asynchrony | Fully asynchronous with partial rollouts at a staleness of 4; partial rollouts keep their original inference log-probabilities and re-prefill after each policy update | | Optimizer | Muown, lr 3e-6, no weight decay or warmup, grad clip 1.0; Muon momentum 0.95 (Nesterov), 10 Newton–Schulz iterations, update scale 0.5; Adam β1 = β2 = 0.95; FP32 master weights and row state carried over from SFT; MoE router frozen | | Task mix (report) | agentic + competitive coding 68%, general tool use 12%, aesthetic design 13%, context following 3%, cybersecurity 4% | | Steps and cost | 30 steps each. Pro: $2.62M, 123.1 h, 75B tokens, 753k samples. Flash: $0.854M, 81.8 h, 81.4B tokens, 753k samples. Total dashboard cost $3,474,715. Pro cost split: rollout 43.8%, training 43.5%, grader 12.7%; Flash 44.9 / 40.9 / 14.2 | | Dashboard start/stop | Pro started 2026-09-15 10:32 UTC, stopped 2026-09-20 (5 d 07:29); Flash started 15:16 UTC, stopped 2026-09-19 (3 d 11:05); streaming since 2026-09-16 04:00 UTC | | Progress | DeepSWE v1.1 avg@3 (mini-swe-agent harness): Pro 58.4 → 72.57, Flash 48.7 → 65.68; training-task pass rate (dynsam/avg@n) Pro 0.633 (+0.068 vs step 1), Flash 0.644 (+0.130) | | Grading | Groupwise Reward Synthesis (offline rubrics; R = R_test · S_sol · S_beh) on high-pass-rate code tasks; Groupwise Advantage Redistribution (online SFT-trained agentic grader ranks passing patches on five dimensions, zeroes confirmed hacks, redistributes positive advantage with a capped common factor); group-relative length penalty; segment-level penalties for format and tool-call errors with conserved advantage mass | | Reward-hacking defense | Mid-training self-correction data; environment preparation (artifact, cache and git cleanup, container-level network isolation); a hack agent iterated until no exploit remains; offline trajectory audits during training; confirmed-hack share below 2% throughout | | Harnesses | Four training mini-harnesses per domain; held-out harness (codex, claude code, mini-swe-agent) mean pass@1 on DeepSWE rose from about 50% to 66%; dashboard step 30 shows 23 trained harnesses over 22,729 rollouts, the largest at 18.9% | ## 4. Infrastructure, component by component (report §6 + dashboard) - **Agent Loop and trajectory hierarchy.** Rollout is agent-centric: each sequence runs an Agent Loop that owns environment setup, interaction, reward evaluation and cleanup, exposes a request endpoint and calls the engine token-in, token-out (only the new suffix is tokenized after prefix matching). Trajectories are Sample → Sequence → Context → Segment; only model-generated segments carry loss. A Penalty Module separates detection (rules: infrastructure failures not attributable to the model, garbled tokens, unavailable tools, repetition) from effect (mask, advantage shaping, monitor, early stop), and escalates: a context with no surviving model turns is dropped, a sequence with no surviving context gets zero advantage, a sample with no surviving sequence is rejected. - **Harness Pool.** Ray actors, but not one per rollout: a dedicated actor per harness instance and Agent Loop would exhaust the GCS node's file descriptors, so fixed-size pools of persistent host actors carry many tenants each (agent loops on the model side sharing one endpoint, inference proxy and tokenizer; harness instances on the environment side sharing one imported codebase). Blocking work runs on background threads. Different harness codebases run in separate pools with a fixed actor budget set at startup. - **Payload Porter (control plane vs data plane).** Each finished sequence's payload (token ids, log-probs, MoE routing ids, top-p candidate indices, multimodal data, gigabytes per trajectory in GUI tasks) is written once to a distributed key-value store (Ray object store or TransferQueue); the driver schedules on metadata only (scalar rewards, per-context lengths, keys). The groupwise grader runs fully asynchronously and rewrites group rewards on return; the sampler accepts or rejects groups by pass rate on metadata; a yield hook packs micro-batches without touching tensors; one packer per training TP group fetches only the rows its CP window touches. Multimodal items are encoded data-parallel across ranks, then redistributed to the ranks holding the tokens. - **Sample Mixer.** Across 25 profiled sources, mean generated tokens vary 90× and active rollout duration 66×. Four mechanisms: adaptive rollout concurrency (budget (1 + p_i)·B_i/r_i with a shared oversampling factor), adaptive rollout scheduling (deficit-corrected weighted round-robin, α = 0.5), predictive rollout dispatch (admit a rollout only if a safety-margin multiple of its estimated KV demand fits the target rank; bound expected inference concurrency by the CUDA-graph max running requests), and sample replay (reuse completed groups from slow sources in the first collection step after startup or recovery, within each source's staleness limit; startup collection otherwise takes about 1.8× as long). - **Training–inference consistency.** Experts are quantize-dequantized to MXFP4 after each update under the Humming kernel's numerics so both engines hold identical weights; Rollout Routing Replay records expert indices at rollout and replays them in training; top-p candidate sets are recorded as a fixed-shape full-vocabulary bitmap (typically fewer than five tokens at top-p 0.97) and training log-probabilities are renormalized within them. A `train_infer_diff/new_infer/kl` metric is streamed on the dashboard. - **Context Cache and DFlash.** Per-dialogue persistent KV plus routing records, candidate sets and visual inputs across turns within a policy version; state lives in HBM during generation and in a pinned host pool during tool time, moved on side CUDA streams. Block-6 FP8 DFlash drafter, initially trained on the SFT policy and finetuned on early RL logs: +31.3% accepted length over the inherited MTP-3, about +6% throughput for block 6 over block 8, about +10.3% per-node throughput for FP8 on the long-context workload. - **Training engine.** Megatron-LM at 1M context; SWA layers exchange only window-sized KV under context parallelism; optimizer states in CPU memory; policy-gradient and OPD losses fused into one kernel. ## 5. Operational evidence (the part RL-Infra-Bench should mine) Report Figure 12 and the dashboard notices agree and complement each other: | Incident | Source | Detail | |---|---|---| | GPU memory double-bit errors | report §5.5; dashboard "vram issue on one node" | primary infrastructure failure class for both runs | | Kubernetes failure in the cyber-task cluster | report; dashboard "infra error on one of datasets was not correctly detected over the past ~3 hours" | Flash restarted from step 15 | | Grader unreachable over the network | report; dashboard | Pro restarted after step 14; Xiaomi also "removed the cyber dataset from the upcoming pro run, since we observed some bad patterns in the rollout logs" | | Inference failures | report | partial-rollout startup bias exhausted both the GPU and pinned host KV pools; one harness produced rollouts less than half as long as the others on the same dataset, skewing dispatch estimates for two steps | | Training failure: GPU OOM from MoE imbalance | report; dashboard "pro run restarted at step 17 due to a GPU OOM issue caused by expert load imbalance" | an EP rank received over 30× the mean token load inside one micro-batch despite a balanced full batch; parallelism adjusted to reduce activation memory | | Driver failure: CPU OOM during packing | report | late in Flash as longer sequences raised per-node data volume beyond host memory, even with packing distributed across nodes | | Router load collapse | report Figure 11 | with a trainable router, over 20 steps CV 0.78 → 2.0, peak load 6× → 16×, cold experts 0.5% → 22%; restoring router weights alone restored balance with unchanged benchmarks; frozen router keeps CV ≈ 0.7, peak ≈ 5.5×, cold ≈ 1% | | Mid-run curriculum change | dashboard | "we filtered out tasks that are relatively easy for the current pro model" | Dashboard batch composition at Pro step 30: 25 sources; code 1,033 prompts (65.8%), visual 305 (19.5%), general 176 (11.3%), chat 54 (3.4%), cyber 0. Dynamic-sampler feed at Pro step 30: accepted 2,208 / 1,568 target, judged 2,155, pass 0.606 (n = 3,644), remaining 332 + 537 partial + 191 rewarding, prewarm 256. Flash step 31 (at stop): accepted 2,704 / 1,568, pass 0.547, with 1,001 partial. The streamed metric tags (2,077 in total) include `partial/avg_staleness`, `dynsam/infra_error/seq_rate`, `env/active`, `timing_s/outer_gen` and `timing_s/trainer_ops`; the charts render on canvas and were not read numerically here. ## 6. Benchmarks (Table 3, baselines at maximum reasoning effort) | Benchmark | Pro | Flash | V2.5-Pro | Opus 5 | GPT-5.6 Sol | Fable 5 | |---|---|---|---|---|---|---| | DeepSWE v1.1 | 71.9 | 67.9 | 19.0 | 74.0 | 73.0 | 70.0 | | ProgramBench | 26.5 | 26.0 | 12.5 | 37.0 | 25.0 | 33.0 | | Terminal-Bench 4.0 | 34.9 | 28.8 | 1.5 | 49.0 | 39.9 | 42.4 | | Terminal-Bench 2.1 | 89.9 | 87.6 | 65.2 | 89.1 | 88.8 | 84.3 | | AutomationBench v1.0.6 | 53.1 | 52.3 | 16.0 | 50.3 | 45.8 | 46.2 | | Toolathlon-Verified | 76.9 | 73.6 | 49.1 | 80.6 | 74.9 | 77.9 | | OSWorld-Verified | 82.0 | 80.8 | — | 83.4 | 83.0 | 86.0 | | CyberGym (corrected environments) | 94.0 | 95.1 | 40.0 | — | — | — | | ExploitBench | 47.9 | 25.3 | 16.6 | 70.0 | 78.5 | 78.0 | Two readings matter for the audit. The 55-point gap between Terminal-Bench 2.1 (saturated, 89.9) and 4.0 (34.9) confirms that the two versions must never share a score series. And the MiMo Code Bench, MiMo Cyber Bench and MiMo Visual Coding numbers are internal and unreproducible. The Distill-9B experiments (Table 6) show RL on the released environments lifting SWE-bench Verified 61.1 → 66.2 and Terminal-Bench 2.1 37.1 → 52.8, and multi-harness RL improving all 21 dataset–harness pairs. ## 7. Discrepancies and open questions 1. Flash size: 310B (report) versus 309B (model card, news). Table 1 prints "42T" active parameters for Pro; the text and card say 42B. 2. Tokens per step: report 2.7–3.7B; blog and news 3.5–3.7B; dashboard step 30 shows 3.43B (Pro) and 3.7B (Flash). 3. Cyber share: the report says 4% of RL tasks; the dashboard shows cyber removed from Pro after the grader incident (0.0% at step 30) while Flash step 31 still accepted 145 cyber groups against a target of 64. 4. Cluster size: "thousands of GPUs" and dollar costs only; no GPU type or count. The Forkast estimate of about $432k per day is derived from the dashboard, not stated by Xiaomi (secondary). 5. Environments: announced as open, not yet published; framework code is. The report's abstract says "we open-source the training dynamics, RL environments, and RL framework"; treat the environments as pending until a dataset appears. 6. Forkast's claim of a "Claude Distill Requests: hidden" dashboard field was not present in the dashboard text read here (secondary, unverified). 7. The report attributes the CyberGym score to "corrected" evaluation environments (footnote 1); the correction method is the report's own sanitizer-report matching, so the number is not comparable to the public CyberGym leaderboard. ## 8. Relevance to the coverage audit and RL-Infra-Bench MiMo-V2.6 becomes the eighteenth documented pipeline on the coverage page and the most detailed first-party account so far of mixed-task asynchronous agentic RL operations: an honest failure ledger with a timeline, a cost decomposition across rollout, training and grader, per-source scheduling formulas, and a streamed metric catalogue. It grounds every RL-Infra-Bench mechanism family with a human incident: - placement and partitioning: EP-rank token imbalance OOM at step 17, fixed by changing parallelism; router freeze after load collapse; - weight synchronization: quantize-dequantize to MXFP4 after each update so rollout and trainer weights match (transport itself not described); - asynchrony and staleness: staleness 4 with partial rollouts that keep their original inference log-probabilities; per-source staleness limits in recovery replay; - train/inference numeric mismatch: R3 routing replay, top-p candidate replay, Humming-consistent MXFP4; - rollout-side throughput: Harness Pool multi-tenancy, predictive dispatch bounded by KV capacity and CUDA-graph concurrency, context cache tiering during tool time; - fault tolerance and recovery: DBEs, a Kubernetes failure, an unreachable grader endpoint, CPU OOM in packing; sample replay after recovery; - throughput engineering: SWA-bounded context parallelism, fused losses, block-6 FP8 DFlash. Candidate source-backed task shapes (each scales down by mechanism, not size): "the grader endpoint becomes unreachable mid-step; the run must stall or fail cleanly without training on unjudged groups"; "one EP rank receives 30× the mean tokens in a micro-batch; diagnose from memory telemetry and fix by repacking or parallelism"; "after a restart the dispatcher's length priors are biased by short rollouts and the KV pools are exhausted"; "router drift produces cold experts; detect via CV, peak load and cold-expert fraction and freeze the router"; "one harness yields rollouts half as long as its peers; the sample mixer must keep per-source shares". The open verl `mimo-oss` branch, uni-agent gateway and MiMoAgent (with Kubernetes and Modal backends and DeepSWE, TerminalBench and ProgramBench adapters) are also a plausible pipeline pack for RL-Infra-Bench once the environments ship.