Gemini 3.1 Pro Preview vs Kimi K3
Gemini 3.1 Pro Preview vs Kimi K3
Independent runs now hand Kimi K3 the agent-work sweep: Toolathlon-Verified 76.5 vs 61.1 and Agents' Last Exam 28.3 vs 16.4 overall pass rate — both real gaps, every number independently run. (An earlier version of this page said Gemini led Agents' Last Exam; that was our metric mix-up — partial-credit score vs pass rate — corrected 2026-08-17.) On Terminal-Bench the gap looks huge (88.3 vs 65.8) but Kimi's side is still its own claim, so keep skepticism there. Gemini 3.1 Pro ties on the reasoning/coding trio — HLE no-tools, GPQA, LiveCodeBench — and costs less ($2/$12 vs $3/$15).
Benchmark by benchmark
| Benchmark | Gemini 3.1 Pro Preview | Kimi K3 | Signal |
|---|---|---|---|
| 🧠 Reasoning(no tools) | 47.0independent | 46.9independent | Tie |
| 🧠 Reasoning(with tools) | 51.4self-reported | 56.0self-reported | Unverified |
| 💻 Terminal ops | 65.8independent | 88.3self-reported | Real gap |
| 🛠️ Long-horizon coding | Not sourced yet | 69.0independent | |
| 🧰 Multi-tool chores | 61.1independent | 76.5independent | Real gap |
| 💼 Professional work | 16.4independent | 28.3independent | Real gap |
| 🐛 Bug fixing | 78.8independent | 93.4independent | Tainted |
| 🔬 Expert science Q&A | 95.5independent | 92.9independent | Tie |
| ⌨️ Contest coding | 88.5independent | 87.2independent | Tie |
In plain English: Scores that use different tool setups aren’t directly comparable — see each row’s signal before reading the ranking.
Who should pick which
- Multi-tool chores and professional agent tasks → Kimi K3 (two real gaps with both sides independently run (toolathlon.xyz, Snorkel) — the confirmed part of the sweep)
- Reasoning and contest coding on a budget → Gemini 3.1 Pro Preview (statistical ties on HLE no-tools, GPQA and LiveCodeBench at a lower list price)
- Terminal-heavy work → Kimi K3 (hedged pick: the 22-point lead is real-sized but Kimi's number is self-reported — if that risk bothers you, the verified runner-up is Opus 4.8's 78.9, not Gemini)
Pricing
| Input / 1M tokens | $2.00 | $3.00 |
| Output / 1M tokens | $12.00 | $15.00 |
FAQ
Is Gemini 3.1 Pro Preview better than Kimi K3?
Independent runs now hand Kimi K3 the agent-work sweep: Toolathlon-Verified 76.5 vs 61.1 and Agents' Last Exam 28.3 vs 16.4 overall pass rate — both real gaps, every number independently run. (An earlier version of this page said Gemini led Agents' Last Exam; that was our metric mix-up — partial-credit score vs pass rate — corrected 2026-08-17.) On Terminal-Bench the gap looks huge (88.3 vs 65.8) but Kimi's side is still its own claim, so keep skepticism there. Gemini 3.1 Pro ties on the reasoning/coding trio — HLE no-tools, GPQA, LiveCodeBench — and costs less ($2/$12 vs $3/$15).
Which is cheaper, Gemini 3.1 Pro Preview or Kimi K3?
Gemini 3.1 Pro Preview costs less per output token ($12.00 vs $15.00 per 1M tokens, official listed rates).