The Model Gap

Claude Opus 4.8 vs Gemini 3.1 Pro Preview

Claude Opus 4.8 vs Gemini 3.1 Pro Preview

Independent runs now give Claude Opus 4.8 three clean real wins on agent work, each confirmed on both sides: Terminal-Bench 2.1 (78.9 vs 65.8), Toolathlon-Verified (76.2 vs 61.1), and Agents' Last Exam (27.0 vs 16.4 overall pass rate). The reasoning and coding picture is all statistical ties — HLE without tools, GPQA Diamond, LiveCodeBench. (An earlier version of this page said Agents' Last Exam favored Gemini; that was our error — we had mixed the board's partial-credit score with its pass rate. Corrected 2026-08-17.) Gemini's remaining case is price: $2/$12 vs $5/$25.

Benchmark by benchmark

BenchmarkClaude Opus 4.8Gemini 3.1 Pro PreviewSignal
🧠 Reasoning(no tools)Tie
🧠 Reasoning(with tools)Unverified
💻 Terminal opsReal gap
🛠️ Long-horizon codingNot sourced yet
🧰 Multi-tool choresReal gap
💼 Professional workReal gap
🐛 Bug fixingTainted
🔬 Expert science Q&ATie
⌨️ Contest codingTie

In plain English: Scores that use different tool setups aren’t directly comparable — see each row’s signal before reading the ranking.

Who should pick which

  • Agentic work — terminal, multi-tool, professional tasks Claude Opus 4.8 (three real gaps, every number on both sides independently run (tbench.ai, toolathlon.xyz, Snorkel))
  • Budget-sensitive reasoning and coding Gemini 3.1 Pro Preview (ties Opus 4.8 on HLE no-tools, GPQA and LiveCodeBench at less than half the price)
  • Want every number verifiable Claude Opus 4.8 (since independent boards covered both models, Opus 4.8's wins are the ones that survived — the earlier 'only Gemini is verified' framing is obsolete)

Pricing

Input / 1M tokens$5.00$2.00
Output / 1M tokens$25.00$12.00

FAQ

Is Claude Opus 4.8 better than Gemini 3.1 Pro Preview?
Independent runs now give Claude Opus 4.8 three clean real wins on agent work, each confirmed on both sides: Terminal-Bench 2.1 (78.9 vs 65.8), Toolathlon-Verified (76.2 vs 61.1), and Agents' Last Exam (27.0 vs 16.4 overall pass rate). The reasoning and coding picture is all statistical ties — HLE without tools, GPQA Diamond, LiveCodeBench. (An earlier version of this page said Agents' Last Exam favored Gemini; that was our error — we had mixed the board's partial-credit score with its pass rate. Corrected 2026-08-17.) Gemini's remaining case is price: $2/$12 vs $5/$25.
Which is cheaper, Claude Opus 4.8 or Gemini 3.1 Pro Preview?
Gemini 3.1 Pro Preview costs less per output token ($12.00 vs $25.00 per 1M tokens, official listed rates).