Claude Opus 4.8 vs Gemini 3.1 Pro Preview
Claude Opus 4.8 vs Gemini 3.1 Pro Preview
Independent runs now give Claude Opus 4.8 three clean real wins on agent work, each confirmed on both sides: Terminal-Bench 2.1 (78.9 vs 65.8), Toolathlon-Verified (76.2 vs 61.1), and Agents' Last Exam (27.0 vs 16.4 overall pass rate). The Terminal-Bench one is the weakest: both rows are tbench.ai's, but with different agent scaffolds (Claude Code for Opus, Gemini CLI for Gemini), and vals.ai's single Terminus 2 harness puts them at 71.91 and 70.79, a tie. The reasoning and coding picture: HLE without tools is a statistical tie. GPQA Diamond and, since 2026-09-29, LiveCodeBench no longer belong in that group — both are graded saturated (grade D), so the close gaps there (92.42 vs 95.45, 87.82 vs 88.48) reflect ceiling-out tests, not evenly-matched models. (An earlier version of this page said Agents' Last Exam favored Gemini; that was our error — we had mixed the board's partial-credit score with its pass rate. Corrected 2026-08-17.) LiveBench adds one more tie (76.2 vs 77.0), and AA-AnalystAgent, added 2026-09-29, another: 45.0 vs 41.25 pass^5, both run by Artificial Analysis, inside that benchmark's 11.2-point band. ARC-AGI-2 again doesn't yield a comparable line, since Opus 4.8's tracked score sits at its High tier (72.1) while Gemini's is a flat, untiered 77.1 — different conditions, no matched tier to compare. HMMT Feb 2026 doesn't add a usable line either: both scores (95.45 vs 94.70) carry MathArena's own disclosed contamination-warning, since every model this site tracks was released after the competition happened, so the whole benchmark is graded tainted rather than read as a real gap or a tie. Gemini's remaining case is price: $2/$12 vs $5/$25.
Claude Opus 4.8 vs Gemini 3.1 Pro Preview: benchmark by benchmark
| Benchmark | Claude Opus 4.8 | Gemini 3.1 Pro Preview | Signal |
|---|---|---|---|
| Reasoning(no tools) | 48.7independent | 47.0independent | Tie |
| Reasoning(with tools) | 57.9self-reported | 51.4self-reported | Unverified |
| Terminal ops | 78.9independent | 65.8independent | Real gap |
| Long-horizon coding | 59.0independent | Not sourced yet | |
| Multi-tool chores | 76.2independent | 61.1independent | Real gap |
| Professional work | 27.0independent | 16.4independent | Real gap |
| Bug fixing | 88.6independent | 78.8independent | Tainted |
| Expert science Q&A | 92.4independent | 95.5independent | Tainted |
| Contest coding | 87.8independent | 88.5independent | Tainted |
| Compositional visual reasoning(high) | 72.1independent | Not sourced yet | |
| Compositional visual reasoning(untiered) | Not sourced yet | 77.1independent | |
| Composite score across 7 domains | 76.2independent | 77.0independent | Tie |
| Competition mathematics | 95.5independent | 94.7independent | Tainted |
| Spreadsheet & document analysis | 45.0independent | 41.3independent | Tie |
Across 14 rows comparing Claude Opus 4.8 and Gemini 3.1 Pro Preview: 3 rows carry a confirmed real gap, 3 rows land inside the noise band, and 8 rows carry other caveats — unverified, tainted, or run on mismatched tool setups. Read each row’s Signal label before ranking Claude Opus 4.8 against Gemini 3.1 Pro Preview.
Who should pick Claude Opus 4.8, and who should pick Gemini 3.1 Pro Preview
- Agentic work — terminal, multi-tool, professional tasks → Claude Opus 4.8 (three real gaps, every number on both sides independently run (tbench.ai, toolathlon.xyz, Snorkel); the Terminal-Bench one mixes agent scaffolds and ties on vals.ai's single harness, so lean on the other two)
- Budget-sensitive reasoning → Gemini 3.1 Pro Preview (ties Opus 4.8 on HLE no-tools (GPQA and LiveCodeBench are now excluded as saturated) at less than half the price)
- Want every number verifiable → Claude Opus 4.8 (since independent boards covered both models, Opus 4.8's wins are the ones that survived — the earlier 'only Gemini is verified' framing is obsolete)
Claude Opus 4.8 vs Gemini 3.1 Pro Preview pricing
| Input / 1M tokens | $5.00 | $2.00 |
| Output / 1M tokens | $25.00 | $12.00 |
Claude Opus 4.8 vs Gemini 3.1 Pro Preview FAQ
Is Claude Opus 4.8 better than Gemini 3.1 Pro Preview?
Independent runs now give Claude Opus 4.8 three clean real wins on agent work, each confirmed on both sides: Terminal-Bench 2.1 (78.9 vs 65.8), Toolathlon-Verified (76.2 vs 61.1), and Agents' Last Exam (27.0 vs 16.4 overall pass rate). The Terminal-Bench one is the weakest: both rows are tbench.ai's, but with different agent scaffolds (Claude Code for Opus, Gemini CLI for Gemini), and vals.ai's single Terminus 2 harness puts them at 71.91 and 70.79, a tie. The reasoning and coding picture: HLE without tools is a statistical tie. GPQA Diamond and, since 2026-09-29, LiveCodeBench no longer belong in that group — both are graded saturated (grade D), so the close gaps there (92.42 vs 95.45, 87.82 vs 88.48) reflect ceiling-out tests, not evenly-matched models. (An earlier version of this page said Agents' Last Exam favored Gemini; that was our error — we had mixed the board's partial-credit score with its pass rate. Corrected 2026-08-17.) LiveBench adds one more tie (76.2 vs 77.0), and AA-AnalystAgent, added 2026-09-29, another: 45.0 vs 41.25 pass^5, both run by Artificial Analysis, inside that benchmark's 11.2-point band. ARC-AGI-2 again doesn't yield a comparable line, since Opus 4.8's tracked score sits at its High tier (72.1) while Gemini's is a flat, untiered 77.1 — different conditions, no matched tier to compare. HMMT Feb 2026 doesn't add a usable line either: both scores (95.45 vs 94.70) carry MathArena's own disclosed contamination-warning, since every model this site tracks was released after the competition happened, so the whole benchmark is graded tainted rather than read as a real gap or a tie. Gemini's remaining case is price: $2/$12 vs $5/$25.
Which is cheaper, Claude Opus 4.8 or Gemini 3.1 Pro Preview?
Gemini 3.1 Pro Preview costs less per output token ($12.00 vs $25.00 per 1M tokens, official listed rates).