MiniMax

MiniMax M3

Four of the five benchmark scores tracked here for MiniMax M3 come from independent evaluators (Vals AI and LiveBench) — but the fifth, Terminal-Bench 2.1, is MiniMax's own 66.0% self-report, and Vals AI's independent bash-only harness measured just 53.56% on the same nominal benchmark, a 12.4-point gap the vendor number doesn't disclose.

MiniMax M3’s 5 benchmark scores on this page were verified against their sources on or after 2026-08-26.

Released
2026-06-01
License
open-weights
Context window
1M tokens
Knowledge cutoff
Not disclosed

The verified record

Against the 77 head-to-head comparisons MiniMax M3 shares with other tracked models: 27 real gaps, 3 inside the noise band, and 47 we will not call.

A gap counts for MiniMax M3 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — MiniMax M3 trails on 26 of them.

LiveBenchComposite score across 7 domains

±2.7 is noise
Behind
Claude Fable 5 15.7 · GPT-5.6 Sol 13.7 · Claude Opus 5 12.8 · Kimi K3 11.9 · Gemini 3.7 Flash 11.5 · Qwen3.8-Max 11.2 · Grok 4.6 10.7 · Muse Spark 1.2 10.7 · DeepSeek V4 Pro (0813) 10.1 · Gemini 3.1 Pro Preview 9.7 · Claude Opus 4.8 8.9 · GLM-5.3 8.8 · Claude Sonnet 5 8.7 · DeepSeek V4 Flash (0731) 6.9 · GPT-5.6 Luna 6.3 · GLM-5.2 5.9

LiveCodeBenchContest coding

±3.1 is noise
Ahead
GLM-5.2 +12.7
Behind
Claude Fable 5 7.6 · Claude Opus 5 6.9 · Gemini 3.7 Flash 6.5 · Gemini 3.1 Pro Preview 6.3 · Grok 4.6 6.0 · Qwen3.8-Max 5.7 · Claude Opus 4.8 5.7 · DeepSeek V4 Pro (0813) 5.4 · DeepSeek V4 Flash (0731) 5.1 · Kimi K3 5.0
Tie
3 models within ±3.1

No verdict for MiniMax M3 anywhere on GPQA Diamond, SWE-bench Verified (saturated); Terminal-Bench 2.1 (nothing independently confirmed on both sides).

  • None of MiniMax M3’s reasoning comparisons are independently confirmed on both sides yet.
  • None of MiniMax M3’s agentic comparisons are independently confirmed on both sides yet.

MiniMax M3 API pricing

$0.30 in / $1.20 out per 1M tokens official pricing

MiniMax M3 benchmark scores

MiniMax M3 benchmark scores, provenance, and source links
BenchmarkScore
GPQA Diamondsaturated[1]
Expert science Q&A — not ranked at any gap size
SWE-bench Verifiedsaturated[2]
Bug fixing — not ranked at any gap size
LiveCodeBench[3]
Contest coding · ±3.1 is noise
Terminal-Bench 2.1[4]
Terminal ops · ±10.6 is noise
LiveBench[5]
Composite score across 7 domains · ±2.7 is noise

Who ran these numbers: 4 of 5 independent — vals.ai (3), livebench.ai (1); vendor self-reported (1).

  1. GPQA Diamond: ±1.44 margin of error; rank 13/135 among all models Vals AI tracks. Value only appears after full client-side hydration -- vals.ai's model card renders this via an animated JS counter that shows 0.0% on first paint to static scrapers; confirmed via direct DOM JavaScript extraction.
  2. SWE-bench Verified: ±1.94 margin; rank 44/86 as of today. Vals AI's own launch post (Jun 2, 2026) ranked this identical 75.00% score #17 -- the rank has drifted purely because 20+ more models were added to the pool since, the score itself hasn't moved. Difficulty-band pass rates: 85% (<15min, 194 tasks), 73% (15m-1h, 261 tasks), 48% (1-4h, 42 tasks), 33% (>4h, 3 tasks) -- weights out to ~75.3%, internally consistent with the headline number.
  3. LiveCodeBench: ±1.05 margin; rank 54/140.
  4. Terminal-Bench 2.1: Vendor's own number: internal infra, 8C16G sandbox, 2hr timeout, 128K max output tokens, Terminus 2 scaffold, temp=1/top_p=0.95. Not used but worth flagging: Vals AI independently measured only 53.56% (rank 39/57) with a bash-only harness -- a 12.4pt vendor-vs-independent gap. MiniMax-M3 has no entry on the official Terminal-Bench 2.1 leaderboard (hub.harborframework.com, 17 total submissions, all frontier proprietary agents) and Artificial Analysis does not expose this sub-score as accessible text anywhere (checked model page DOM, raw HTML/JSON, and main leaderboard table).
  5. LiveBench: Reported as LiveBench's "Global Average" figure on the 2026-06-25 release. Row label on site: "Minimax M3". Subscores: Reasoning 74.5, Coding 68.2, Agentic Coding 40.7, Mathematics 76.9, Data Analysis 76.2, Language 76.8, Instruction Following 57.5. Cost per successful task: $0.060.

Notes on the record

MiniMax M3 is a 428B-total / 23B-active-parameter mixture-of-experts model — confirmed via Artificial Analysis's embedded model JSON, where activeParams (23) + passiveParams (405) sum to 428, as of 2026-08-26. It launched June 1, 2026 per MiniMax's own blog and Artificial Analysis's internal releaseDate field; Vals AI lists May 31, 2026 instead, a one-day gap most plausibly explained by China-time (UTC+8, the blog's timezone) versus US-Pacific-time (Vals AI's likely logging timezone) rather than a genuine dispute over ship date. It carries a 1,000,000-token context window and ships under MiniMax's own "Community License" (Hugging Face license id: minimax-community) — a Llama-style restricted-commercial-use open-weight license, not an OSI-approved license like MIT or Apache. No knowledge cutoff date is disclosed anywhere checked: not MiniMax's blog, not Artificial Analysis, not the Hugging Face model card (all as of 2026-08-26).

Pricing on MiniMax's own docs (platform.minimax.io/docs/guides/pricing-paygo) is framed as a "Permanent 50% off" promotion: $0.30 in / $1.20 out per million tokens for inputs up to 512K tokens, rising to $0.60 in / $2.40 out above that threshold. List (non-promotional) prices for the same two tiers are $0.60/$2.40 and $1.20/$4.80 respectively, plus a separate 1.5x priority-service multiplier not reflected in the base rates above. The cheapest independently verified third-party host, checked live via OpenRouter, is CoreWeave at $0.23 in / $0.96 out per million tokens — roughly 20-25% below MiniMax's own price.

Five benchmark scores are tracked for this page. Four are independently measured: GPQA Diamond 92.68% (Vals AI, ±1.44 margin, rank 13/135, as of 2026-08-26 — a figure that required a workaround, since Vals AI's model card renders scores via an animated JS counter that shows 0.0% to static scrapers on first paint, and only surfaced via direct DOM JavaScript execution after full client hydration); SWE-bench Verified 75.00% (Vals AI, ±1.94 margin, as of 2026-08-26); LiveCodeBench 82.15% (Vals AI, ±1.05 margin, rank 54/140, as of 2026-08-26); and LiveBench's Global Average 67.3% (LiveBench-2026-06-25 release, as of 2026-08-26; subscores: Reasoning 74.5, Coding 68.2, Agentic Coding 40.7, Mathematics 76.9, Data Analysis 76.2, Language 76.8, Instruction Following 57.5; cost per successful task $0.060). The SWE-bench Verified number is a "score stable, rank isn't" case: Vals AI's own June 2, 2026 launch post ranked this identical 75.00% score #17; the same score today (2026-08-26) sits at #44/86 purely because 20+ more models have since been added to the tracked pool, not because the score itself changed.

The fifth score, Terminal-Bench 2.1, is this page's one vendor self-report: MiniMax's blog claims 66.0% (internal infra, 8C16G sandbox, 2-hour timeout, 128K max output tokens, Terminus 2 scaffold, temp=1/top_p=0.95, as of 2026-08-26). Vals AI's independent bash-only harness measured just 53.56% (rank 39/57) on the same nominal benchmark — a 12.4-point gap. MiniMax-M3 has no submission on the official Terminal-Bench 2.1 leaderboard (hub.harborframework.com, 17 total entries, all frontier proprietary agents), so no third source exists to arbitrate between the two numbers. Artificial Analysis does not expose Terminal-Bench or HLE sub-scores as accessible text for this model anywhere checked — model page DOM, raw HTML/JSON, and the main comparison table were all inspected; the main comparison table shows only a composite Intelligence Index of 45 and aggregate cost/speed figures, while the model page itself renders evals as unlabeled sparkline charts with no numeric text at all. Separately, Hugging Face shows about 205,085 downloads of MiniMaxAI/MiniMax-M3 in the past month (as of 2026-08-26), which reads as sustained open-weight adoption rather than launch-week novelty.

Compare with

FAQ

Is MiniMax M3 good for coding?

Coding results are mixed on the independently measured benchmarks. LiveCodeBench looks solid at 82.15% (rank 54/140, upper third) under Vals AI's testing, but SWE-bench Verified's 75.00% now sits at rank 44/86 — essentially median, even though the score itself hasn't moved since a stronger #17 ranking in June (the field has simply grown). LiveBench's Agentic Coding subscore is only 40.7, its weakest category by a wide margin, suggesting more complex, multi-step coding workflows are less proven than the model's other coding results.

Are MiniMax M3's benchmark numbers independently verified?

Mostly. Four of the five benchmarks tracked here — GPQA Diamond, SWE-bench Verified, LiveCodeBench, and LiveBench — come from independent evaluators (Vals AI and LiveBench.ai), not MiniMax. The exception is Terminal-Bench 2.1: MiniMax's self-reported score runs 12.4 points above what Vals AI measured independently on the same benchmark, and no third source exists to arbitrate which number is right.

How much does MiniMax M3 cost through the API?

MiniMax's own pricing for inputs up to 512K tokens is $0.30 per million input tokens and $1.20 per million output tokens, framed by the company as a permanent half-off promotional rate rather than its list price. Rates double for longer inputs, and third-party host CoreWeave has been found undercutting MiniMax's own rate by roughly a fifth to a quarter.

Is MiniMax M3 open source?

It's open-weight, not open-source in the OSI sense. MiniMax releases the weights under its own "Community License" (Hugging Face id: minimax-community), a Llama-style license that restricts certain commercial uses — not an OSI-approved license like MIT or Apache. Hugging Face shows about 205,085 downloads of the model in the past month, indicating meaningful real-world adoption despite the license restrictions.

What context window does MiniMax M3 support, and does it have a published knowledge cutoff?

MiniMax M3 supports a 1 million token context window. No knowledge cutoff date has been disclosed, however — it doesn't appear on MiniMax's blog, on Artificial Analysis, or on the model's Hugging Face card, so this site cannot state one.

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.