Xiaomi's MiMo team announced MiMo-V2.6: two open-weight omnimodal models, Pro and Flash, plus an UltraSpeed inference mode for Pro. The launch thread leads with a specific claim: "Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks." The technical report's own comparison table, and Artificial Analysis's independent read of the model posted the next day, are both worth checking against that sentence.
"On par" bundles two different claims, and only one of them holds
Table 3 in the technical report lists 17 benchmarks with Opus 5 and Sol scores reported for 14 and 15 of them respectively. Tallying MiMo-V2.6-Pro against each name separately, rather than against the bundled "Opus 5 and Sol":
Against Claude Opus 5: Pro wins 3 (AutomationBench, Terminal Bench 2.1, MiMo Visual Coding), ties 1 (Agents' Last Exam, 31.6 apiece), and loses 10 — including a 35-point gap on GDPval-AA 2.1, a 14.1-point gap on Terminal Bench 4.0, and a 22.1-point gap on ExploitBench (47.9 vs 70.0). That's not parity by any reasonable reading of a 3–10–1 record.
Against GPT-5.6 Sol: Pro wins 8 and loses 7, including a 30.6-point gap on ExploitBench — genuinely close to a coin flip. The "on par with Sol" half of the claim holds up on Xiaomi's own numbers; the "on par with Opus 5" half doesn't.
The comparison set is also one generation behind Xiaomi's own marketing page
Opus 5 and Sol aren't the current frontier. Claude Fable 5.1 and GPT-6 Astra both launched three weeks before this release and both superseded the models in Table 3 — neither appears in the technical report's comparison table. They do show up, selectively, on Xiaomi's marketing page, and where they do the picture gets worse for MiMo-Pro: Terminal Bench 4.0 (Astra 59.6, Fable 5.1 55.1, Pro 34.9), ExploitGym (Astra 42.4 vs Pro 17.8), and MiMo Visual Coding — Xiaomi's own in-house benchmark (Astra 82.2, Fable 5.1 74.4, Pro 72.3). Pro does beat Astra on MiMo Code Bench, AutomationBench, and GDPval-AA 2.1, so it isn't a clean sweep either way — but the categories that most resemble real agentic execution are exactly where the two current frontier models pull furthest ahead, and exactly the two models the technical report's table leaves out.
The DeepSWE number is missing the one disclosure that would make it checkable
DeepSeek's own V4.1 Flash model card ran the identical benchmark — DeepSWE v1.1 — across eight scaffolds and found an 8.7-point spread, with the number DeepSeek used to claim it "edged past" Opus 5 turning out to be the single highest-scoring scaffold of the eight. MiMo-V2.6-Pro's 71.9 on the same benchmark comes with no scaffold disclosed, which matters more here than for most labs: ASI-Bench already measured MiMo-V2.5-Pro's own score moving from 16.17 to 23.25 just from switching harnesses. A model already shown scaffold-sensitive, on a benchmark already shown to swing 8.7 points by scaffold, reported without saying which one produced the number, isn't currently checkable the way DeepSeek's own disclosure was.
The RSI framing has no hedge at all
The launch post's second sentence — "This marks a key step in our exploration of the RSI path: scaling RL compute on verifiable, complex tasks, so the model can continuously expand its capability frontier through exploration and feedback" — is the entire treatment recursive self-improvement gets. No safety language, no boundary-setting, no "we're early" qualifier. OpenAI's RSI report states it doesn't "yet know how to safely get all the way to aligned, full RSI." Z.ai's Infra Agent essay argues explicitly that "choosing objectives, setting boundaries, and assessing risk remain human responsibilities." Even Meta's AIRA₃ announcement, the least-hedged of the three, opens with "we're early, and hard problems are still ahead of us" before reaching for the term. Among four labs now on record using it this month, Xiaomi is the only one with zero accompanying safety language.
UltraSpeed's mechanism is real; its quality claim isn't shown
The "up to 20x faster... at the same quality" claim for UltraSpeed isn't a black box — Xiaomi has documented it as DFlash speculative decoding (block-level masked parallel token prediction) combined with selective FP4 quantization of MoE expert matrices, keeping routers and attention at higher precision. That's a genuine, named mechanism. What's missing is a benchmark row for UltraSpeed itself next to standard Pro — quantizing MoE weights to 4-bit is a well-studied source of measurable quality loss even with quantization-aware training, and "at the same quality" isn't verified anywhere in what's published.
The RL cost figure is the scaling run, not the model
Xiaomi reports the RL phase completed 30 steps over ~750,000 trajectories in under six days, at about 2.62M for Pro — a combined $3.47M, already circulating in coverage headlined around that number. Worth being precise: this is the disclosed RL-scaling phase specifically, not pretraining and not the base checkpoints' underlying compute. Flash's run cost less than a third of Pro's and produced the larger relative gain, the expected shape for a model with more room to improve from a lower starting point. Streaming the run live across six days is a transparency practice this blog hasn't seen from another lab this year, worth noting as a genuine positive even without a way to confirm how continuous the stream actually was.
The MOF-materials co-design case and the Lean 4 formalization of Li and Yorke's theorem that Xiaomi highlights alongside the launch are both single worked examples, not benchmarked evaluations — covered in more detail separately, including how little of the MOF claim has actually been physically verified.
Where it actually sits: Artificial Analysis's own numbers
Artificial Analysis's own charts, posted the day after launch.
AA's bar chart puts MiMo-V2.6-Pro at 46 on the Intelligence Index (v4.3) — every model scoring above it (Fable 5.1, Astra, Opus 5, Muse Spark 1.3, Sol) is closed, and the two closest open-weight rivals, Qwen3.8 Max and Kimi K3, sit one and two points behind. "Strongest open-source model to date" checks out against AA's own numbers, not just Xiaomi's. AA also states the model is 1.02 trillion total parameters, 42 billion active, and prices it at 0.87 per million tokens — a Pareto-frontier position on AA's own cost-versus-intelligence plot, which carries more weight than most self-drawn Pareto claims precisely because a third party plotted it.
The index also explains the Opus 5 gap above. Three of its ten components (GDPval-AA, AutomationBench, Terminal Bench 4.0) closely resemble rows in Table 3 where Pro lost badly to Opus 5 — yet the full ten-part composite gap is a contained 5 points (46 vs 51), not the wider spread those specific losses alone would suggest. The other seven components are doing real work narrowing it. MiMo-V2.6-Pro's actual position isn't "on par with Opus 5" or "far behind" — it's a model that loses ground on a subset of agentic-execution tasks and makes most of it back elsewhere.
One loose end: a second, differently-styled version of this frontier chart is circulating, crediting every point to "Artificial Analysis public benchmark data" except MiMo-V2.6-Pro's, which it sources separately to "Xiaomi MiMo." Taken literally that would mean AA hasn't independently run the model and is relaying Xiaomi's own number. Against that, this blog's own Muse Spark 1.3 coverage already established that AA marks a genuinely unverified score "not currently available" rather than publish a number — and MiMo-V2.6-Pro's AA chart entry got a specific bar, not that flag. The more likely reading is AA's own figure is real; the second chart's footnote, from a source that can't be identified here, is a complication worth flagging rather than resolving.
MiMo-V2.5 shipped under an MIT license; if V2.6 follows that precedent — Xiaomi's own materials don't state terms explicitly — it joins the 81% of large Chinese open-weight releases this blog has already found carrying Apache or MIT terms.