Claude Opus 4.8 vs Gemini 3.1 Pro Preview
Claude Opus 4.8 vs Gemini 3.1 Pro Preview
Independent runs now give Claude Opus 4.8 three clean real wins on agent work, each confirmed on both sides: Terminal-Bench 2.1 (78.9 vs 65.8), Toolathlon-Verified (76.2 vs 61.1), and Agents' Last Exam (27.0 vs 16.4 overall pass rate). The reasoning and coding picture is all statistical ties — HLE without tools, GPQA Diamond, LiveCodeBench. (An earlier version of this page said Agents' Last Exam favored Gemini; that was our error — we had mixed the board's partial-credit score with its pass rate. Corrected 2026-08-17.) Gemini's remaining case is price: $2/$12 vs $5/$25.
Benchmark by benchmark
| Benchmark | Claude Opus 4.8 | Gemini 3.1 Pro Preview | Signal |
|---|---|---|---|
| 🧠 Reasoning(no tools) | 48.7independent | 47.0independent | Tie |
| 🧠 Reasoning(with tools) | 57.9self-reported | 51.4self-reported | Unverified |
| 💻 Terminal ops | 78.9independent | 65.8independent | Real gap |
| 🛠️ Long-horizon coding | 59.0independent | Not sourced yet | |
| 🧰 Multi-tool chores | 76.2independent | 61.1independent | Real gap |
| 💼 Professional work | 27.0independent | 16.4independent | Real gap |
| 🐛 Bug fixing | 88.6independent | 78.8independent | Tainted |
| 🔬 Expert science Q&A | 92.4independent | 95.5independent | Tie |
| ⌨️ Contest coding | 87.8independent | 88.5independent | Tie |
In plain English: Scores that use different tool setups aren’t directly comparable — see each row’s signal before reading the ranking.
Who should pick which
- Agentic work — terminal, multi-tool, professional tasks → Claude Opus 4.8 (three real gaps, every number on both sides independently run (tbench.ai, toolathlon.xyz, Snorkel))
- Budget-sensitive reasoning and coding → Gemini 3.1 Pro Preview (ties Opus 4.8 on HLE no-tools, GPQA and LiveCodeBench at less than half the price)
- Want every number verifiable → Claude Opus 4.8 (since independent boards covered both models, Opus 4.8's wins are the ones that survived — the earlier 'only Gemini is verified' framing is obsolete)
Pricing
| Input / 1M tokens | $5.00 | $2.00 |
| Output / 1M tokens | $25.00 | $12.00 |
FAQ
Is Claude Opus 4.8 better than Gemini 3.1 Pro Preview?
Independent runs now give Claude Opus 4.8 three clean real wins on agent work, each confirmed on both sides: Terminal-Bench 2.1 (78.9 vs 65.8), Toolathlon-Verified (76.2 vs 61.1), and Agents' Last Exam (27.0 vs 16.4 overall pass rate). The reasoning and coding picture is all statistical ties — HLE without tools, GPQA Diamond, LiveCodeBench. (An earlier version of this page said Agents' Last Exam favored Gemini; that was our error — we had mixed the board's partial-credit score with its pass rate. Corrected 2026-08-17.) Gemini's remaining case is price: $2/$12 vs $5/$25.
Which is cheaper, Claude Opus 4.8 or Gemini 3.1 Pro Preview?
Gemini 3.1 Pro Preview costs less per output token ($12.00 vs $25.00 per 1M tokens, official listed rates).