Moonshot AI

Kimi K3

Twelve of Kimi K3's thirteen tracked scores are independently confirmed; the exception is the with-tools HLE score, still Moonshot's own number. Its Terminal-Bench 2.1 row is now vals.ai's independent run, 80.9, 7.4 points under the 88.3 Moonshot reported.

Kimi K3 benchmarks and pricing, every number sourced: 13 tracked Kimi K3 benchmark scores (12 independently run, 1 still resting on a vendor’s own claim), priced at $3.00 per million input tokens and $15.00 per million output.

Kimi K3 architecture: Natively multimodal Mixture-of-Experts; 2.8T total parameters (104B activated per token); 1M-token context window.

Kimi K3’s 13 benchmark scores on this page were each verified against their sources between 2026-07-16 and 2026-10-01.

Released
2026-07-16
License
open-weights
Context window
1M tokens
Knowledge cutoff
Not disclosed
Parameters
2.8T (104B active)
Architecture
Natively multimodal Mixture-of-Experts

Kimi K3’s verified record

Kimi K3’s featured comparison is Gemini 3.1 Pro Preview: 2 leads, 3 ties, and 6 not callable. Kimi K3 is priced at $3.00/$15.00 per 1M tokens (in/out) vs Gemini 3.1 Pro Preview’s $2.00/$12.00. Full verdict →

Against the 242 head-to-head comparisons Kimi K3 shares with other tracked models: 64 real gaps, 57 inside the noise band, and 121 we will not call.

A gap counts for Kimi K3 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Kimi K3 trails on 26 of them.

HLE · no tools Reasoning

±2 is noise
Behind
MiMo-V2.6-Pro −2.5 · GPT-5.6 Sol −2.6 · GPT-6.1 Sol −6.0 · GPT-6 Astra −7.8 · Claude Opus 5 −8.0 · Claude Sonnet 5.5 −8.1 · Claude Fable 5.1 −12.2 · Claude Opus 5.5 −14.5 · 1 superseded: Claude Fable 5 −8.6
Ahead
Qwen3.8-Max +3.8 · Grok 4.7 +3.8 · Grok 4.6 +4.0 · GLM-5.3 +4.6 · Claude Sonnet 5 +5.6 · DeepSeek V4 Pro (0813) +5.9 · GLM-5.3-Flash +7.0 · GPT-5.6 Luna +7.4 · GPT-6 Luna +8.4 · Qwen3.8-Flash-Next +8.9 · 2 superseded: GLM-5.2 +5.8 · DeepSeek V4 Flash (0731) +8.3
Tie
8 models within ±2
Unverified
2 models — vendor-reported on one side

LiveBench Composite score across 7 domains

±2.7 is noise
Behind
Claude Opus 5.5 −4.0 · Claude Fable 5.1 −4.2 · 1 superseded: Claude Fable 5 −3.8
Ahead
Qwen3.8-Flash-Next +3.0 · GLM-5.3 +3.1 · Claude Sonnet 5 +3.2 · GPT-5.6 Luna +5.6 · GPT-6 Luna +7.2 · GLM-5.3-Flash +7.6 · MiniMax M3 +11.9 · 3 superseded: Claude Opus 4.8 +3.0 · DeepSeek V4 Flash (0731) +5.0 · GLM-5.2 +6.0
Tie
13 models within ±2.7

Agents' Last Exam Professional work

±3.2 is noise
Behind
Claude Opus 5 −3.9 · GPT-6 Sol −3.9 · Muse Spark 1.3 −3.9 · GPT-6 Astra −5.9 · Claude Opus 5.5 −9.9
Ahead
GPT-6 Luna +3.3 · Gemini 3.1 Pro Preview +11.9 · 1 superseded: GLM-5.2 +7.9
Tie
5 models within ±3.2
Unverified
6 models — vendor-reported on one side

DeepSWE Long-horizon coding

±9.5 is noise
Ahead
Qwen3.8-Max +12.0 · Claude Sonnet 5 +15.0 · 4 superseded: Claude Opus 4.8 +10.0 · Muse Spark 1.2 +14.0 · DeepSeek V4 Flash (0731) +16.0 · GLM-5.2 +25.0
Tie
10 models within ±9.5
Unverified
10 models — vendor-reported on one side
Setup-dependent
2 models — scored on a different harness

AnalystAgent Spreadsheet & document analysis

±11.2 is noise
Behind
GPT-6 Astra −12.5 · Claude Opus 5 −15.0 · Claude Fable 5.1 −18.8 · 1 superseded: Gemini 3.7 Flash −21.3
Ahead
Tie
6 models within ±11.2

ARC-AGI-2 · max Compositional visual reasoning

±9.2 is noise
Behind
Claude Fable 5.1 −29.6 · Claude Opus 5 −30.0 · GPT-5.6 Sol −32.1 · GPT-6 Astra −34.6 · 1 superseded: Claude Fable 5 −28.8
Tie
4 models within ±9.2

Terminal-Bench 2.1 Terminal ops

±10.6 is noise
Ahead
MiMo-V2.6-Pro +13.1 · Tencent Hy4 preview +25.8 · MiniMax M3 +27.3 · 1 superseded: GLM-5.2 +13.1
Tie
7 models within ±10.6
Unverified
1 model — vendor-reported on one side
Setup-dependent
19 models — scored on a different harness

Toolathlon-Verified Multi-tool chores

±9.7 is noise
Ahead
Gemini 3.1 Pro Preview +15.4 · 1 superseded: GLM-5.2 +16.6
Tie
4 models within ±9.7
Unverified
10 models — vendor-reported on one side

No verdict for Kimi K3 anywhere on GPQA Diamond, LiveCodeBench, SWE-bench Verified (saturated); HLE · with tools (nothing independently confirmed on both sides); HMMT (contaminated).

Kimi K3 API pricing

$3.00 in / $15.00 out per 1M tokens — official pricing

What Kimi K3 costs per job

Kimi K3 cost for three reference workloads, computed from its list rates
WorkloadTokens in / outCost
One long chat turn100K / 10K$0.450
A codebase review1,000K / 100K$4.50
A day of agent work10,000K / 1,000K$45.00

Computed from Kimi K3’s list rates above — cache discounts and batch tiers are not applied.

Kimi K3 benchmark scores

Kimi K3 benchmark scores, provenance, and source links
BenchmarkScore
HLE(no tools)[1]
Reasoning · ±2 is noise
HLE(with tools)
Reasoning · ±2 is noise
GPQA Diamondsaturated[2]
Expert science Q&A — not ranked at any gap size
Terminal-Bench 2.1[3]
Terminal ops · ±10.6 is noise
DeepSWE[4]
Long-horizon coding · ±9.5 is noise
Toolathlon-Verified[5]
Multi-tool chores · ±9.7 is noise
Agents' Last Exam[6]
Professional work · ±3.2 is noise
SWE-bench Verifiedsaturated[7]
Bug fixing — not ranked at any gap size
LiveCodeBenchsaturated[8]
Contest coding — not ranked at any gap size
ARC-AGI-2(max)[9]
Compositional visual reasoning · ±9.2 is noise
LiveBench[10]
Composite score across 7 domains · ±2.7 is noise
HMMTcontaminated[11]
Competition mathematics — not ranked at any gap size
AnalystAgent[12]
Spreadsheet & document analysis · ±11.2 is noise

Who ran these numbers: 12 of 13 independent — vals.ai (4), artificialanalysis.ai (2), deepswe.datacurve.ai (1), toolathlon.xyz (1), snorkel.ai (1), arcprize.org (1), livebench.ai (1), matharena.ai (1); vendor self-reported (1).

  1. HLE: AA's own run, 'Kimi K3 (max)' (text-only, no tools). Replaces Moonshot's self-reported 43.5.
  2. GPQA Diamond: vals.ai run, rank 10/133 (updated 2026-08-15). Replaces Moonshot's self-reported 93.5.
  3. Terminal-Bench 2.1: vals.ai's own Terminus 2 run, board row 'Kimi K3' (80.90 ± 0.65, no effort parameter listed, pass@1), from its archived Terminal-Bench 2.1 table. Replaces Moonshot's self-reported 88.3 (Kimi Code harness, GitHub README), 7.4 points higher.
  4. DeepSWE: Independent score (69% ± 5%, $4.65/task); vendor self-reported 67.5 is close and consistent.
  5. Toolathlon-Verified: Rank 1 on the independent leaderboard; exact match with vendor self-reported figure.
  6. Agents' Last Exam: Overall pass rate 28.3, board rank 3 when first recorded on 2026-08-17 (Kimi Code, Max; score 51.6; $606). Metric verified correct in the 2026-08-17 normalization.
  7. SWE-bench Verified: vals.ai run, rank 6/83, bash-only harness, $0.76/test (updated 2026-08-14).
  8. LiveCodeBench: vals.ai run, rank 16/138 (updated 2026-08-15).
  9. ARC-AGI-2: Kimi K3's official ARC-AGI-2 leaderboard row, dated 2026-07-16 on arcprize.org. The top of three populated tiers (Max/High/Low).
  10. LiveBench: Board row "Kimi K3" on the 2026-06-25 LiveBench release.
  11. HMMT: MathArena's leaderboard row is "Kimi K3 (Think)", 96.97% across all 33 HMMT Feb 2026 problems (live-verified 2026-08-25 at matharena.ai/?comp=hmmt--hmmt_feb_2026). Carries MathArena's own contamination-warning flag — "Model was released after competition release" — a disclosed risk that applies to every one of this site's 4 tracked HMMT Feb 2026 rows, Kimi K3 included, since all 4 tracked models were released after HMMT Feb 2026 took place.
  12. AnalystAgent: AA's own run, board row 'Kimi K3 (max)'; the board rounds to one decimal. pass@1 57.75, pass@5 72.5.

Notes on the record

Cache-hit input price is $0.30/1M, far below the $3.00 cache-miss price shown (confirmed current on platform.kimi.ai/docs/pricing/chat-k3 as of 2026-08-20, no future price-change date published).

Open weights under a bespoke license, not plain MIT/Apache — commercial use above certain revenue/MAU thresholds requires a separate Moonshot agreement or interface attribution; see the FAQ below for the exact thresholds (source: huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE).

Chat/API access was announced 2026-07-16 (the date tracked as this model's release), but the full 2.8T-parameter weight files weren't posted for public download until 2026-07-26 — a day ahead of the 2026-07-27 date Moonshot had previously targeted (source: mlq.ai/news/moonshot-releases-156-tb-kimi-k3-weights-under-a-custom-commercial-license).

Compare with

FAQ

Has Kimi K3 been independently benchmarked?

Mostly yes. Of the thirteen Kimi K3 benchmark results tracked on this site, twelve come from independent evaluators: GPQA Diamond and SWE-bench Verified and LiveCodeBench (all vals.ai, observed 2026-08-17), Terminal-Bench 2.1 (vals.ai, observed 2026-10-01), DeepSWE (deepswe.datacurve.ai, observed 2026-08-13), Toolathlon-Verified (toolathlon.xyz, observed 2026-07-16), Agents' Last Exam (snorkel.ai, observed 2026-08-17), the no-tools HLE run and AA-AnalystAgent (Artificial Analysis, observed 2026-08-17 and 2026-09-29), ARC-AGI-2 and LiveBench (ARC Prize and LiveBench's own board, observed 2026-08-24), and HMMT Feb 2026 (MathArena, observed 2026-08-25). The remaining one, the with-tools HLE score from Moonshot's own README (observed 2026-08-17), is self-reported. The same README claims 88.3 on Terminal-Bench 2.1; vals.ai's independent run measured 80.9.

What license does Kimi K3 use, and does it restrict commercial use?

Kimi K3 ships under a bespoke "Kimi K3 License," not standard MIT or Apache 2.0. Per the license text (huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE), most commercial use, fine-tuning, and redistribution is permitted, but two thresholds apply: a Model-as-a-Service operator whose aggregate revenue (with affiliates) exceeds $20M over any trailing 12 months must sign a separate agreement with Moonshot before commercial use, and any commercial product built on Kimi K3 that reaches over 100 million monthly active users or $20M/month in revenue must display "Kimi K3" on its user interface. Internal-only use and access through Moonshot's own API or certified partners are exempt from both.

Why does Kimi K3 sometimes cost less than its listed $3 per million input tokens?

The $3.00/1M figure shown on this page is Kimi K3's cache-miss input price. Moonshot's own pricing docs (platform.kimi.ai/docs/pricing/chat-k3) list a separate cache-hit input price of $0.30/1M for tokens that match a recent prompt prefix — a 10x difference — while output pricing stays $15.00/1M either way. No future change date for these rates is published as of 2026-08-20.

When did Kimi K3's open weights actually become downloadable?

Kimi K3's chat and API access opened on 2026-07-16, which is the date tracked as this model's release on this site. The full 2.8-trillion-parameter weight files weren't posted for public download until 2026-07-26 — a day ahead of the 2026-07-27 target Moonshot had previously communicated (source: mlq.ai/news/moonshot-releases-156-tb-kimi-k3-weights-under-a-custom-commercial-license). Anyone evaluating Kimi K3 for self-hosting rather than API use should use the later date.

How does Kimi K3 compare to Gemini 3.1 Pro Preview?

This site runs a dedicated head-to-head verdict comparing Kimi K3 and Gemini 3.1 Pro Preview rather than declaring a single winner in this FAQ. See the related verdict linked on this page for the category-by-category breakdown, since the two models' scores come from different mixes of independent and self-reported sources that don't reduce cleanly to one number.

Further reading

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.