MiniMax
MiniMax M3
All nine benchmark scores tracked here for MiniMax M3 come from independent evaluators (Vals AI, LiveBench and Artificial Analysis) — HLE without tools joined on 2026-10-03 (Artificial Analysis, 39.0). The last to switch was Terminal-Bench 2.1: Vals AI's Terminus 2 run measured 53.56%, 12.4 points under the 66.0% MiniMax reported from its own infrastructure, and replaced that self-report here on 2026-10-01.
MiniMax M3 benchmarks and pricing, every number sourced: 9 tracked MiniMax M3 benchmark scores (9 independently run, 0 still resting on a vendor’s own claim), priced at $0.30 per million input tokens and $1.20 per million output.
MiniMax M3 architecture: Sparse Mixture-of-Experts, natively multimodal; ~428B total parameters (~23B activated per token); 1M-token context window.
- Released
- 2026-06-01
- License
- open-weights
- Context window
- 1M tokens
- Knowledge cutoff
- Not disclosed
- Verified
- sources checked 2026-08-26–2026-10-08
- Parameters
- ~428B (~23B active)
- Architecture
- Sparse Mixture-of-Experts, natively multimodal
MiniMax M3’s verified record
MiniMax M3’s most-compared rival is Qwen3.8-Max: 4 trails and 4 not callable across their 8 shared comparisons. MiniMax M3 is priced at $0.30/$1.20 per 1M tokens (in/out) vs Qwen3.8-Max’s $2.00/$6.00.
Against the 176 head-to-head comparisons MiniMax M3 shares with other tracked models: 87 real gaps, 8 inside the noise band, and 81 we will not call.
A gap counts for MiniMax M3 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — MiniMax M3 trails on 87 of them.
LiveBench Composite score across 7 domains
±2.7 is noiseHLE · no tools Reasoning
±2 is noiseAnalystAgent Spreadsheet & document analysis
±11.2 is noiseTerminal-Bench 2.1 Terminal ops
±10.6 is noiseTerminal-Bench 4.0 Terminal ops
±12.4 is noiseOSWorld 2.0 · v2026 06 24 full Long-horizon computer use
±9.7 is noiseNo verdict for MiniMax M3 anywhere on GPQA Diamond, LiveCodeBench, SWE-bench Verified ( saturated).
- None of MiniMax M3’s coding comparisons are independently confirmed on both sides yet.
MiniMax M3 API pricing
$0.30 in / $1.20 out per 1M tokens — official pricing source
What MiniMax M3 costs per job
| Workload | Tokens in / out | Cost |
|---|---|---|
| One long chat turn | 100K / 10K | $0.042 |
| A codebase review | 1,000K / 100K | $0.420 |
| A day of agent work | 10,000K / 1,000K | $4.20 |
Computed from MiniMax M3’s list rates above — cache discounts and batch tiers are not applied.
MiniMax M3 benchmark scores
| Benchmark | Score |
|---|---|
| GPQA Diamondsaturated[1] Expert science Q&A — not ranked at any gap size | |
| SWE-bench Verifiedsaturated[2] Bug fixing — not ranked at any gap size | |
| LiveCodeBenchsaturated[3] Contest coding — not ranked at any gap size | |
| Terminal-Bench 2.1[4] Terminal ops · ±10.6 is noise | |
| LiveBench[5] Composite score across 7 domains · ±2.7 is noise | |
| AnalystAgent[6] Spreadsheet & document analysis · ±11.2 is noise | |
| HLE(no tools)[7] Reasoning · ±2 is noise | |
| Terminal-Bench 4.0[8] Terminal ops · ±12.4 is noise | |
| OSWorld 2.0(v2026 06_24_full)[9] Long-horizon computer use · ±9.7 is noise |
Who ran these numbers: 9 of 9 independent — vals.ai (5), artificialanalysis.ai (2), livebench.ai (1), osworld-v2.xlang.ai (1).
- GPQA Diamond: ±1.44 margin of error; rank 13/135 among all models Vals AI tracks. Value only appears after full client-side hydration -- vals.ai's model card renders this via an animated JS counter that shows 0.0% on first paint to static scrapers; confirmed via direct DOM JavaScript extraction.
- SWE-bench Verified: ±1.94 margin; rank 44/86 as of today. Vals AI's own launch post (Jun 2, 2026) ranked this identical 75.00% score #17 -- the rank has drifted purely because 20+ more models were added to the pool since, the score itself hasn't moved. Difficulty-band pass rates: 85% (<15min, 194 tasks), 73% (15m-1h, 261 tasks), 48% (1-4h, 42 tasks), 33% (>4h, 3 tasks) -- weights out to ~75.3%, internally consistent with the headline number.
- LiveCodeBench: ±1.05 margin; rank 54/140.
- Terminal-Bench 2.1: vals.ai's own Terminus 2 run, board row 'MiniMax-M3' (53.56 ± 0.75, no effort parameter listed, pass@1), from its archived Terminal-Bench 2.1 table. Replaces MiniMax's self-reported 66 (Terminus 2 scaffold on MiniMax's own infrastructure, 2-hour timeout), 12.4 points higher.
- LiveBench: Reported as LiveBench's "Global Average" figure on the 2026-06-25 release. Row label on site: "Minimax M3". Subscores: Reasoning 74.5, Coding 68.2, Agentic Coding 40.7, Mathematics 76.9, Data Analysis 76.2, Language 76.8, Instruction Following 57.5. Cost per successful task: $0.060.
- AnalystAgent: AA's own run, board row 'MiniMax-M3'. pass@1 44.0, pass@5 73.75 — the widest reliability gap on the tracked set: it gets 44% of attempts right and solves 73.75% of questions at least once, but almost never five times in a row.
- HLE: AA's own run, board row 'MiniMax-M3' (no effort tier published) — text-only, no tools; underlying value 38.97, board-displayed 39.0. Previously this model had no independent HLE row: its GPQA/LiveCodeBench/SWE-bench numbers are vals.ai runs on benchmarks this site grades saturated.
- Terminal-Bench 4.0: vals.ai board row — mini-swe-agent harness (single bash tool), pass@1 averaged over 3 full passes, raw value 1.01 at no effort tier stated. Artificial Analysis same-conditions run (no tier suffix): 2.02.
- OSWorld 2.0: Authors' own run, original task set. Partial reward 22.3, binary accuracy 4.6. Reasoning enabled (no effort tier given by the source), standard tool calls, 500 steps. Shorter budgets on the same release: 8.2 partial at 150 steps, 16.6 at 300. No effort tier is recorded because the board lists the run as reasoning-enabled rather than tiered. Verified twice on 2026-10-08 (direct fetch + reader proxy). Official board 500-step row (the board default; 150/300-step runs read 8.2/16.6); 2026-10-08.
Notes on the record
MiniMax M3 is a 428B-total / 23B-active-parameter mixture-of-experts model — confirmed via Artificial Analysis's embedded model JSON, where activeParams (23) + passiveParams (405) sum to 428, as of 2026-08-26. It launched June 1, 2026 per MiniMax's own blog and Artificial Analysis's internal releaseDate field; Vals AI lists May 31, 2026 instead, a one-day gap most plausibly explained by China-time (UTC+8, the blog's timezone) versus US-Pacific-time (Vals AI's likely logging timezone) rather than a genuine dispute over ship date. It carries a 1,000,000-token context window and ships under MiniMax's own "Community License" (Hugging Face license id: minimax-community) — a Llama-style restricted-commercial-use open-weight license, not an OSI-approved license like MIT or Apache. No knowledge cutoff date is disclosed anywhere checked: not MiniMax's blog, not Artificial Analysis, not the Hugging Face model card (all as of 2026-08-26).
Pricing on MiniMax's own docs (platform.minimax.io/docs/guides/pricing-paygo) is framed as a "Permanent 50% off" promotion: $0.30 in / $1.20 out per million tokens for inputs up to 512K tokens, rising to $0.60 in / $2.40 out above that threshold. List (non-promotional) prices for the same two tiers are $0.60/$2.40 and $1.20/$4.80 respectively, plus a separate 1.5x priority-service multiplier not reflected in the base rates above. The cheapest independently verified third-party host, checked live via OpenRouter, is CoreWeave at $0.23 in / $0.96 out per million tokens — roughly 20-25% below MiniMax's own price.
Nine benchmark scores are tracked for this page, all independently measured. Five of the original set: GPQA Diamond 92.68% (Vals AI, ±1.44 margin, rank 13/135, as of 2026-08-26 — read from the model card's DOM after client hydration, since its animated counter shows 0.0% to static scrapers); SWE-bench Verified 75.00% (Vals AI, ±1.94 margin, as of 2026-08-26); LiveCodeBench 82.15% (Vals AI, ±1.05 margin, rank 54/140, as of 2026-08-26); and LiveBench's Global Average 67.3% (LiveBench-2026-06-25 release, as of 2026-08-26; subscores: Reasoning 74.5, Coding 68.2, Agentic Coding 40.7, Mathematics 76.9, Data Analysis 76.2, Language 76.8, Instruction Following 57.5; cost per successful task $0.060); and AA-AnalystAgent 10.0 pass^5 (Artificial Analysis, as of 2026-09-29). The SWE-bench Verified number is a "score stable, rank isn't" case: Vals AI's own June 2, 2026 launch post ranked this identical 75.00% score #17; the same score today (2026-08-26) sits at #44/86 purely because 20+ more models have since been added to the tracked pool, not because the score itself changed.
The sixth score, Terminal-Bench 2.1, is Vals AI's own Terminus 2 run, 53.56% (±0.75, no effort parameter listed, as of 2026-10-01). It replaced MiniMax's blog claim of 66.0% (internal infra, 8C16G sandbox, 2-hour timeout, 128K max output tokens, Terminus 2 scaffold, temp=1/top_p=0.95). The gap between the two is 12.4 points. MiniMax-M3 has no submission on tbench.ai's official Terminal-Bench 2.1 leaderboard, and vals.ai has since archived its own board. Artificial Analysis does not expose Terminal-Bench or HLE sub-scores as accessible text for this model anywhere checked — model page DOM, raw HTML/JSON, and the main comparison table were all inspected; the main comparison table shows only a composite Intelligence Index of 45 and aggregate cost/speed figures, while the model page itself renders evals as unlabeled sparkline charts with no numeric text at all. Separately, Hugging Face shows about 205,085 downloads of MiniMaxAI/MiniMax-M3 in the past month (as of 2026-08-26), which reads as sustained open-weight adoption rather than launch-week novelty.
Compare with
FAQ
Is MiniMax M3 good for coding?
Coding results are mixed on the independently measured benchmarks. LiveCodeBench looks solid at 82.15% (rank 54/140, upper two-fifths) under Vals AI's testing, but SWE-bench Verified's 75.00% now sits at rank 44/86 — essentially median, even though the score itself hasn't moved since a stronger #17 ranking in June (the field has simply grown). LiveBench's Agentic Coding subscore is only 40.7, its weakest category by a wide margin, suggesting more complex, multi-step coding workflows are less proven than the model's other coding results. Terminal-Bench 2.1 points the same way: 53.56% in Vals AI's independent run.
Are MiniMax M3's benchmark numbers independently verified?
Yes. All nine benchmarks tracked here — GPQA Diamond, SWE-bench Verified, LiveCodeBench, Terminal-Bench 2.1, LiveBench, AA-AnalystAgent, and HLE without tools — come from independent evaluators (Vals AI, LiveBench.ai and Artificial Analysis), not MiniMax. Terminal-Bench 2.1 was the last to switch: MiniMax's self-reported 66.0% runs 12.4 points above Vals AI's independent 53.56%, and this site now records the Vals AI run.
How much does MiniMax M3 cost through the API?
MiniMax's own pricing for inputs up to 512K tokens is $0.30 per million input tokens and $1.20 per million output tokens, framed by the company as a permanent half-off promotional rate rather than its list price. Rates double for longer inputs, and third-party host CoreWeave has been found undercutting MiniMax's own rate by roughly a fifth to a quarter.
Is MiniMax M3 open source?
It's open-weight, not open-source in the OSI sense. MiniMax releases the weights under its own "Community License" (Hugging Face id: minimax-community), a Llama-style license that restricts certain commercial uses — not an OSI-approved license like MIT or Apache. Hugging Face shows about 205,085 downloads of the model in the past month, indicating meaningful real-world adoption despite the license restrictions.
What context window does MiniMax M3 support, and does it have a published knowledge cutoff?
MiniMax M3 supports a 1 million token context window. No knowledge cutoff date has been disclosed, however — it doesn't appear on MiniMax's blog, on Artificial Analysis, or on the model's Hugging Face card, so this site cannot state one.
What is MiniMax M3's strongest ranked result?
LiveBench at 67.3 (2026-08-26); Terminal-Bench 2.1 reads 53.56 well behind it. The numbers above them on this page — GPQA Diamond 92.68 and LiveCodeBench 82.15 — were posted on boards since retired as saturated, so they carry no rank.
Further reading
- GPQA Diamond leaderboard 2026 — MiniMax M3 is one of the 23 models it compares.
- Models with 10M token context windows 2026 — MiniMax M3 is one of the 37 models it compares.
Benchmark guides
- Terminal-Bench 2.1 — what MiniMax M3’s agentic score on it does and doesn’t prove.