News / Model Watch · Open-Weights Release
Xiaomi Open-Sources MiMo-V2.6: 46 on the AA Index, Prices Held at V2.5, the Whole RL Kitchen Sink
Xiaomi just made the strongest open-weights argument of the season: MiMo-V2.6 shipped September 22 as a three-model family — Pro, Flash and Pro-UltraSpeed — with weights and a technical report on Hugging Face, API prices held at V2.5 levels, and a 7,000+ environment RL training stack that went open too. MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index, overtaking Kimi K3 and Qwen3.8 Max as the strongest open-weights model alive. Here's what shipped, what it costs, and what the RL openness means for builders.
MiMo-V2.6 is open and on Xiaomi's open platform API (Sep 22): model names mimo-v2.6-pro, mimo-v2.6-flash, mimo-v2.6-pro-ultraspeed; pricing held at V2.5 levels (V2.5 Pro benchmark: $3/1M output flat, 1M-token context); weights + technical report open on Hugging Face (XiaomiMiMo/mimo-v26 collection).
The headline: 46 on the AA Intelligence Index — past Kimi K3 and Qwen3.8 Max, the strongest open-weights model as of release; still behind closed leaders Claude Fable 5.1 and GPT-6 Astra. Pro-UltraSpeed keeps Pro performance at up to 20x inference speed. Not yet on third-party relays like this one — GLM, DeepSeek, Qwen and Kimi remain the relay shelf today.
What shipped (Sep 22, official channels)
- MiMo-V2.6-Pro. The flagship reasoning model — natively omni-modal, trillion-parameter, aimed at complex long-horizon projects, cybersecurity and research workloads. Call it as
mimo-v2.6-pro. - MiMo-V2.6-Flash. Omni-modal, high-intelligence, low-cost — positioned for high-frequency professional use. Xiaomi's own framing: it fully surpasses the previous MiMo-V2.5-Pro, which our cheapest-Chinese-LLM guide still lists at the V2.5 benchmark.
- MiMo-V2.6-Pro-UltraSpeed. Pro performance with up to 20x inference speed, for latency-sensitive production — available on the open platform API and inside MiMo Desktop.
- Pricing held at V2.5 levels. Xiaomi claims a domestic price-performance record: at equal intelligence, 1/20 to 1/60 the price of overseas models. The V2.5 reference point: $3/1M output flat at 1M-token context.
- Desktop + subscriptions. MiMo Desktop (official release) and paid subscriptions shipped the same day; the ecosystem now spans MiMo Claw (Kingsoft office integration), MiMo Code and MiMo Studio.
Context for the V2.6 line: V2.5 series with AA Index traction → RL training live-streamed (Aug–Sep) → V2.6 open release + Desktop (Sep 22). The RL-as-content playbook is now Xiaomi's signature.
The training story: scaled RL in the open
Per the official post, MiMo-V2.6's RL training is likely the largest RL compute ever sunk into a Chinese open-weights model — and the whole run was streamed live. Under 6 days of Live RL training: Flash and Pro cost roughly $0.85M and $2.62M respectively, each completing 30 steps with ~750k cumulative trajectories; average task pass rates improved 25% and 12%. On the out-of-sample DeepSWE v1.1 long-horizon benchmark, scores rose ~17 points (48.8 → 65.7) for Flash and ~14 points (58.4 → 72.6) for Pro. Training ran at 1M-token context with 3.5–3.7B tokens per step, mixing Code, General, Visual and Cyber task families across multiple harnesses.
On most agent benchmarks, MiMo-V2.6-Pro lands close to Claude Opus 5 and GPT-5.6 Sol. The official demos ("Vibe World") go further: 3D open-world game generation, Blender modeling, embodied manipulation of a Franka Panda arm, Computer-Use workflows, a PFAS-absorbing MOF material screening run, and a full Lean 4 formalization of the Li–Yorke "period three implies chaos" theorem (6,000+ kernel-verified lines).
What's in the open-source package
- Weights + technical report. Hugging Face
XiaomiMiMo/mimo-v26collection — all three tiers, plus MiMo-V2.6-Distill-Qwen-9B. - 7k+ RL task environments. Software engineering, vulnerability reproduction, knowledge work and web design — the environments the family was trained on.
- End-to-end RL training framework. Built on verl / uni-agent / mini-swe-agent — the stack that ran the live streams.
- Minimal composable mini-harnesses. System prompts, tools and context management decoupled; Multi-Harness Training targets cross-framework generalization.
Compare the openness pattern with DeepSeek Vision-Exp MIT weights — the open-weights lane is getting crowded, but nobody else ships the RL kitchen sink with it.
Pricing position vs the field (per 1M tokens)
| Model | Output | Notes |
|---|---|---|
| MiMo V2.6 Pro (Xiaomi open platform) | held at V2.5 | V2.5 benchmark $3 flat · 1M context; names all-lowercase mimo-v2.6-* |
| DeepSeek V4.1 Flash beta | ~$0.63 off-peak | native multimodal beta; coverage |
| Kimi K3 | — | overtaken by V2.6 Pro on AA Index 46 |
| Qwen3.8 Max | — | also passed at 46; Qwen on the relay |
| Overseas frontier (closed) | $10–50 band | Xiaomi: equal-intelligence V2.6 at 1/20–1/60 the price |
MiMo is not on third-party relays yet (as of Sep 25) — table reflects Xiaomi's official open-platform API; relay-shelf models keep their own margins, see the full price guide.
Why the RL openness is the story
The multi-harness recipe — training one model across many agent frameworks and open-sourcing the harnesses themselves — signals that model vendors now optimize directly for the harness layer, not just chat quality. For builders, that means agent-native behavior (tool use, context discipline, multi-step reliability) is becoming a first-class release criterion, not a downstream adaptation. The MiMo ecosystem already runs this logic end-to-end: MiMo Claw pairs flagship models with the Kingsoft office suite as a subscription; MiMo was also wired into agent clients like Hermes Agent with a limited-time free window. The intersection of frontier open models and agent harnesses keeps getting busier.
Primary sources
FAQ (2026)
Is MiMo-V2.6 open source?
Yes. Full family (Pro / Flash / Pro-UltraSpeed) open on Hugging Face (XiaomiMiMo/mimo-v26) with technical report, Sep 22. The package also ships MiMo-V2.6-Distill-Qwen-9B, 7k+ RL task environments, an end-to-end RL framework (verl / uni-agent / mini-swe-agent) and composable mini-harnesses.
What are the API model names and prices?
All-lowercase mimo-v2.6-pro / mimo-v2.6-flash / mimo-v2.6-pro-ultraspeed on Xiaomi's open platform. Prices held at V2.5 levels — the V2.5 Pro benchmark is $3/1M output flat at 1M-token context; Xiaomi frames it as 1/20–1/60 the price of equal-intelligence overseas models.
How good is MiMo-V2.6-Pro really?
46 on the AA Intelligence Index — past Kimi K3 and Qwen3.8 Max, the strongest open-weights model as of release; still behind closed leaders (Claude Fable 5.1, GPT-6 Astra). Live-RL numbers: Flash/Pro cost ~$0.85M/$2.62M in under 6 days, lifting DeepSWE v1.1 from 48.8→65.7 and 58.4→72.6.
What is Pro-UltraSpeed for?
Latency-sensitive production: Pro-level performance with up to 20x inference speed. Available on the open-platform API and inside MiMo Desktop (which shipped the same day).
Can I call MiMo-V2.6 through ChinaModelAPI?
Not yet (as of Sep 25). MiMo is not on the current relay shelf — GLM, DeepSeek, Qwen, Kimi and more are. This page tracks Xiaomi's official API; when MiMo joins the relay we'll update here and the homepage model grid.
Why does the RL openness matter?
7k+ open task environments + training framework + harnesses = the largest RL compute openly documented for a Chinese open-weights model. Vendors optimizing for the agent-harness layer is the signal; MiMo Claw / Code / Studio and the Hermes Agent integration are the ecosystem already running on it.