Xiaomi

MiMo-V2.6-Pro

Xiaomi's new open-weight flagship: a 1.02T-parameter sparse MoE with MIT weights, 1M context and API pricing of $0.435/$0.87 per million tokens. One of its five tracked scores is independent — Artificial Analysis's Humanity's Last Exam run, 49.35 — while the four coding and agent numbers come from Xiaomi's model card, and none of those benchmarks' own boards had listed it when checked.

MiMo-V2.6-Pro benchmarks and pricing, every number sourced: 5 tracked MiMo-V2.6-Pro benchmark scores (1 independently run, 4 still resting on a vendor’s own claim), priced at $0.43 per million input tokens and $0.87 per million output.

MiMo-V2.6-Pro is Mixture-of-Experts with 1.02T total parameters (42B activated per token) and a 1M-token context window.

MiMo-V2.6-Pro’s 5 benchmark scores on this page were each verified against their sources between 2026-09-22 and 2026-09-23.

Released
2026-09-22
License
open-weights
Context window
1M tokens
Knowledge cutoff
Not disclosed
Parameters
1.02T (42B active)
Architecture
Mixture-of-Experts

MiMo-V2.6-Pro’s verified record

MiMo-V2.6-Pro’s most-compared rival is DeepSeek V4 Pro (0813): 1 lead and 4 not callable across their 5 shared comparisons. MiMo-V2.6-Pro is priced at $0.43/$0.87 per 1M tokens (in/out) vs DeepSeek V4 Pro (0813)’s $1.32/$3.96.

Against the 106 head-to-head comparisons MiMo-V2.6-Pro shares with other tracked models: 21 real gaps, 6 inside the noise band, and 79 we will not call.

A gap counts for MiMo-V2.6-Pro only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — MiMo-V2.6-Pro trails on 5 of them.

HLE · no tools Reasoning

±2 is noise
Behind
GPT-6 Astra 5.3 · Claude Opus 5 5.5 · Claude Fable 5.1 9.8 · Claude Opus 5.5 12.0 · 1 superseded: Claude Fable 5 6.1
Ahead
Gemini 3.1 Pro Preview +2.4 · Kimi K3 +2.5 · Step 5 Preview +2.9 · Grok 4.7 +6.3 · Qwen3.8-Max +6.4 · Grok 4.6 +6.5 · GLM-5.3 +7.1 · Claude Sonnet 5 +8.1 · DeepSeek V4 Pro (0813) +8.4 · GLM-5.3-Flash +9.5 · GPT-5.6 Luna +9.9 · GPT-6 Luna +10.9 · Qwen3.8-Flash-Next +11.4 · 3 superseded: Muse Spark 1.2 +3.9 · GLM-5.2 +8.3 · DeepSeek V4 Flash (0731) +10.8
Tie
6 models within ±2
Unverified
1 model — vendor-reported on one side

No verdict for MiMo-V2.6-Pro anywhere on Agents' Last Exam, DeepSWE, Terminal-Bench 2.1, Toolathlon-Verified (nothing independently confirmed on both sides).

  • None of MiMo-V2.6-Pro’s coding comparisons are independently confirmed on both sides yet.
  • None of MiMo-V2.6-Pro’s agentic comparisons are independently confirmed on both sides yet.

MiMo-V2.6-Pro API pricing

$0.43 in / $0.87 out per 1M tokens official pricing

What MiMo-V2.6-Pro costs per job

MiMo-V2.6-Pro cost for three reference workloads, computed from its list rates
WorkloadTokens in / outCost
One long chat turn100K / 10K$0.052
A codebase review1,000K / 100K$0.522
A day of agent work10,000K / 1,000K$5.22

Computed from MiMo-V2.6-Pro’s list rates above — cache discounts and batch tiers are not applied.

MiMo-V2.6-Pro is one of 2 Xiaomi models tracked on this site, at these official list prices.

Xiaomi model pricing, official list rates
ModelIn / 1MOut / 1M
MiMo-V2.6-Flash$0.14$0.28
MiMo-V2.6-Pro$0.43$0.87

MiMo-V2.6-Pro benchmark scores

MiMo-V2.6-Pro benchmark scores, provenance, and source links
BenchmarkScore
Terminal-Bench 2.1[1]
Terminal ops · ±10.6 is noise
DeepSWE[2]
Long-horizon coding · ±9.5 is noise
Toolathlon-Verified[3]
Multi-tool chores · ±9.7 is noise
Agents' Last Exam[4]
Professional work · ±3.2 is noise
HLE(no tools)[5]
Reasoning · ±2 is noise

Who ran these numbers: 1 of 5 independent — artificialanalysis.ai (1); vendor self-reported (4).

  1. Terminal-Bench 2.1: Xiaomi's model-card evaluation table reports 89.9 with no effort tier stated. tbench.ai's 2.1 board did not list this model when checked 2026-09-22.
  2. DeepSWE: Xiaomi's model-card evaluation table reports 71.9. Xiaomi's announcement instead reports 72.6 as this RL run's after-training endpoint on the same held-out benchmark (up from 58.4 before it) — two official numbers, conflict flagged, neither independently confirmed; deepswe.datacurve.ai did not list this model when checked 2026-09-22.
  3. Toolathlon-Verified: Xiaomi's model-card evaluation table reports 76.9 for the Pro. toolathlon.xyz's board listed only MiMo V2.5 when checked 2026-09-22.
  4. Agents' Last Exam: Xiaomi's model-card evaluation table reports 31.6 for the Pro, split not stated. snorkel.ai's board listed only MiMo V2.5 when checked 2026-09-22.
  5. HLE: Artificial Analysis's own run (text-only 2,158-question subset; the chart data reads 0.4935). AA lists the model as "MiMo-V2.6-Pro" with no effort level, so which reasoning tier produced the score is not stated. The model page on AA shows only the Intelligence Index, not this sub-score.

Notes on the record

Pricing checked 2026-09-22 against Xiaomi's official model page (mimo.mi.com/models/en-US/mimo-v2.6-pro): $0.435 per million input tokens, $0.87 output, $0.0036 cached input — one flat rate with no length threshold, time-of-day discount or promotional condition in the docs. Xiaomi states the V2.6 prices are unchanged from the V2.5 series (announcement page, 2026-09-22).

Spec from the HuggingFace model card (XiaomiMiMo/MiMo-V2.6-Pro-RL, MIT license, uploaded 2026-09-21): sparse MoE, 1.02T total / 42B activated parameters, 1M-token context, 128K max output, text/image/video/audio input. Knowledge cutoff is not disclosed. The official announcement's Update Time reads September 22 while Artificial Analysis and the HuggingFace upload date the release September 21; this site records the official page's date.

Provenance: one tracked score is independent — Artificial Analysis's Humanity's Last Exam run, 49.35 (text-only, no effort tier shown), listed on AA's HLE chart though not on its model page, which shows only the Intelligence Index of 46 (corrected 2026-09-23; this page first said no tracked benchmark had an independent score). The four coding and agent scores are Xiaomi's own model-card figures. Checked live 2026-09-22: tbench.ai, deepswe.datacurve.ai, toolathlon.xyz, snorkel.ai, livebench.ai, arcprize.org and matharena.ai do not list the model; a vals.ai mirror likewise showed only the MiMo V2.5 generation. Xiaomi's two official sources disagree on DeepSWE v1.1: the model card's evaluation table says 71.9, while the announcement's training write-up reports 72.6 as the after-training endpoint of this RL run (up from 58.4 before it), on the same held-out benchmark used to argue the training generalizes — a different kind of claim from the card's headline eval, but still Xiaomi's own number either way. The card figure is recorded, the conflict flagged, neither independently confirmed.

Compare with

FAQ

Has MiMo-V2.6-Pro been independently benchmarked?

Partially. Artificial Analysis runs it independently: besides a composite Intelligence Index of 46, its Humanity's Last Exam chart lists the model at 49.35 (text-only, no effort tier shown; checked 2026-09-23) — the only independent score on this page. tbench.ai, deepswe.datacurve.ai, toolathlon.xyz, snorkel.ai, livebench.ai, arcprize.org and matharena.ai did not list the model when checked on 2026-09-22, so the four coding and agent scores below are Xiaomi's own.

How much does the MiMo-V2.6-Pro API cost?

$0.435 per million input tokens and $0.87 per million output tokens on Xiaomi's official pricing page (checked 2026-09-22), with cached input at $0.0036. It is one flat rate — no length threshold, time-of-day discount or promotion — and Xiaomi says it matches the V2.5-series price.

Is MiMo-V2.6-Pro open source?

Yes. The weights are downloadable from HuggingFace (XiaomiMiMo/MiMo-V2.6-Pro-RL) under an MIT license, released together with the technical report, RL environments and training framework per Xiaomi's announcement on 2026-09-22.

How good is MiMo-V2.6-Pro at coding and agent tasks?

Xiaomi's own table reports Terminal-Bench 2.1 at 89.9, DeepSWE v1.1 at 71.9, Toolathlon-Verified at 76.9 and Agents' Last Exam at 31.6. Every one is a model-card figure with no independent confirmation yet, and Xiaomi's announcement separately reports 72.6 for the same benchmark while describing this RL run's training curve — treat the coding claims as vendor-reported until a board re-runs them.

Why do two different DeepSWE scores circulate for MiMo-V2.6-Pro?

Xiaomi's HuggingFace model card lists 71.9 in its evaluation table. Its announcement instead reports 72.6 while describing an RL training run: DeepSWE v1.1 improved from 58.4 to 72.6 over the run, offered as evidence the training generalizes beyond its own task mix. Both are Xiaomi's own numbers for what should be the same held-out benchmark. This site records the model-card figure (71.9) and flags the conflict; neither has been independently confirmed (checked 2026-09-22).

Further reading

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.