Anthropic
Claude Opus 5
Anthropic pitches Opus 5 as "close to the frontier intelligence of Claude Fable 5 at half the price" — but that price is just Claude Opus 4.8's unchanged $5/$25 rate carried over, not a cut. What's real is the scorecard: all ten benchmark results tracked here are independent runs, not Anthropic's own launch-day figures.
Claude Opus 5 benchmarks and pricing, every number sourced: 10 tracked Claude Opus 5 benchmark scores (10 independently run, 0 still resting on a vendor’s own claim), priced at $5.00 per million input tokens and $25.00 per million output.
Claude Opus 5’s 10 benchmark scores on this page were each verified against their sources between 2026-08-17 and 2026-10-01.
Version history: succeeded Claude Opus 4.8 (2026-05-28).
- Released
- 2026-07-24
- License
- proprietary
- Context window
- 1M tokens
- Knowledge cutoff
- 2026-05
- Parameters
- Not disclosed
- Architecture
- Not disclosed
Claude Opus 5’s verified record
Claude Opus 5’s featured comparison is Claude Sonnet 5: 3 leads, 1 tie, and 4 not callable. Claude Opus 5 is priced at $5.00/$25.00 per 1M tokens (in/out) vs Claude Sonnet 5’s $2.00/$10.00. Full verdict →
Against the 214 head-to-head comparisons Claude Opus 5 shares with other tracked models: 65 real gaps, 53 inside the noise band, and 96 we will not call.
A gap counts for Claude Opus 5 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Claude Opus 5 trails on 7 of them.
HLE · no tools Reasoning
±2 is noiseLiveBench Composite score across 7 domains
±2.7 is noiseAgents' Last Exam Professional work
±3.2 is noiseDeepSWE Long-horizon coding
±9.5 is noiseARC-AGI-2 · max Compositional visual reasoning
±9.2 is noiseAnalystAgent Spreadsheet & document analysis
±11.2 is noiseNo verdict for Claude Opus 5 anywhere on Terminal-Bench 2.1 (every independently confirmed comparison inside the noise band); GPQA Diamond, LiveCodeBench, SWE-bench Verified (saturated).
What changed from Claude Opus 4.8 to Claude Opus 5
The 9 benchmarks both models have been scored on, using the same variant each time. A raw Claude Opus 5 gain is not a real gain until it clears that benchmark’s own noise band, so each row below carries the verdict and not just the arithmetic.
| Benchmark | Claude Opus 4.8 | Claude Opus 5 | Change | Verdict |
|---|---|---|---|---|
| HLE | 48.7 | 54.9 | +6.2 | Real gap |
| Terminal-Bench 2.1 | 78.9 | 89.1 | +10.2 | Setup-dependent |
| DeepSWE | 59 | 74 | +15.0 | Real gap |
| Agents' Last Exam | 27 | 32.2 | +5.2 | Real gap |
| SWE-bench Verified | 88.6 | 97 | +8.4 | Tainted |
| GPQA Diamond | 92.42 | 93.43 | +1.0 | Tainted |
| LiveCodeBench | 87.82 | 89.03 | +1.2 | Tainted |
| LiveBench | 76.2 | 80.1 | +3.9 | Real gap |
| AnalystAgent | 45 | 53.75 | +8.8 | Tie |
Claude Opus 5 API pricing
$5.00 in / $25.00 out per 1M tokens — official pricing
What Claude Opus 5 costs per job
| Workload | Tokens in / out | Cost |
|---|---|---|
| One long chat turn | 100K / 10K | $0.750 |
| A codebase review | 1,000K / 100K | $7.50 |
| A day of agent work | 10,000K / 1,000K | $75.00 |
Computed from Claude Opus 5’s list rates above — cache discounts and batch tiers are not applied.
Claude Opus 5 is one of 7 Anthropic models tracked on this site, at these official list prices.
| Model | In / 1M | Out / 1M |
|---|---|---|
| Claude Sonnet 5.5 | $2.00 | $10.00 |
| Claude Opus 5.5 | $4.00 | $20.00 |
| Claude Fable 5.1 | $10.00 | $50.00 |
| Claude Opus 5 | $5.00 | $25.00 |
| Claude Sonnet 5 | $2.00 | $10.00 |
| Claude Fable 5(superseded) | $10.00 | $50.00 |
| Claude Opus 4.8(superseded) | $5.00 | $25.00 |
Claude Opus 5 benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise | |
| Terminal-Bench 2.1[2] Terminal ops · ±10.6 is noise | |
| DeepSWE[3] Long-horizon coding · ±9.5 is noise | |
| SWE-bench Verifiedsaturated[4] Bug fixing — not ranked at any gap size | |
| LiveCodeBenchsaturated[5] Contest coding — not ranked at any gap size | |
| GPQA Diamondsaturated[6] Expert science Q&A — not ranked at any gap size | |
| Agents' Last Exam[7] Professional work · ±3.2 is noise | |
| ARC-AGI-2(max)[8] Compositional visual reasoning · ±9.2 is noise | |
| LiveBench[9] Composite score across 7 domains · ±2.7 is noise | |
| AnalystAgent[10] Spreadsheet & document analysis · ±11.2 is noise |
Who ran these numbers: 10 of 10 independent — artificialanalysis.ai (3), vals.ai (3), deepswe.datacurve.ai (1), snorkel.ai (1), arcprize.org (1), livebench.ai (1).
- HLE: AA live leaderboard value, 'Claude Opus 5 (Adaptive Reasoning, Max Effort)' — text-only, no tools. Re-verified on the board 2026-08-19. Corrected 2026-08-19: this row used to cite AA's Opus 5 launch article, which reports 53.0 — the 54.9 was always the board's number, so the citation pointed at a page that did not support it.
- Terminal-Bench 2.1: AA's own run, chart-JSON label 'Claude Opus 5 (max)'; page prose gives the fuller label 'Claude Opus 5 (Adaptive Reasoning, Max Effort)' for the same score. A lower-effort 'Claude Opus 5 (xhigh)' variant also appears in the same chart at 88.0. Not on tbench.ai's official board, which lists only older Anthropic models (Fable 5, Opus 4.8, Opus 4.7, Sonnet 5). Cross-check (2026-10-01): vals.ai's archived Terminus 2 table lists it at 84.64, with Opus 4.8 as a refusal fallback on 9 of 267 tasks; counting those as failures gives 81.27.
- DeepSWE: 74%±4 pass rate on the mini-swe-agent harness (all models run identically); board updated 2026-08-13.
- SWE-bench Verified: vals.ai run, rank 1 of 83 (updated 2026-08-14); Anthropic self-reports 96.0. Benchmark near-saturated — five models at 95%+.
- LiveCodeBench: vals.ai, updated 2026-08-15; re-verified on the vals.ai model page 2026-08-17.
- GPQA Diamond: vals.ai run, rank 7/135 (board updated 2026-08-19).
- Agents' Last Exam: Overall pass rate 32.2 (Claude Code, High, the board's best effort for this model; score 55.9; est. cost $1,108). Per-effort results on the board: High 32.2, Max 30.9, XHigh 30.3, Medium 28.9, Low 27.6. Snorkel's data notes record that the board's upstream source re-published every Claude Opus 5 figure on 2026-09-24, moving its Overall best tier (High) from 31.6 to 32.2. This site's earlier 27.0 (Max; score 49.0; est. cost $7,239, recorded 2026-08-20) no longer appears; the current Max run reads 30.9 (score 52.7).
- ARC-AGI-2: Claude Opus 5's official ARC-AGI-2 leaderboard row, dated 2026-07-24 on arcprize.org. Higher than its only other populated tier, High (88.3%).
- LiveBench: Board row "Claude 5 Opus Thinking Max Effort (LiveBench's own row label reverses this site's word order; corroborated as the same model by a real GitHub issue on livebench/livebench)" on the 2026-06-25 LiveBench release.
- AnalystAgent: AA's own run, board row 'Claude Opus 5 (Adaptive Reasoning, Max Effort)'; the board rounds this to 53.8. pass@1 63.75, pass@5 73.75.
Notes on the record
Successor to Claude Opus 4.8 at the same $5/$25 price (confirmed current on Anthropic's pricing/model overview page as of 2026-08-20). Configurable 'effort' toggle (shared machinery, first shipped on Claude Opus 4.7) — Anthropic's docs list five levels (low, medium, high, xhigh, max), defaulting to high on the Claude API and Claude Code (source: platform.claude.com/docs/en/about-claude/models/overview). Anthropic positions it as 'close to the frontier intelligence of Claude Fable 5 at half the price'.
Max output 128k (300k via batch beta, using the `output-300k-2026-03-24` beta header on the Message Batches API, per the same source). Also available via Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry.
Compare with
FAQ
Has Claude Opus 5 been independently benchmarked?
Yes, on every score this site currently tracks for it. All ten results — Humanity's Last Exam (no tools), Terminal-Bench 2.1, DeepSWE, SWE-bench Verified, LiveCodeBench, GPQA Diamond, Agents' Last Exam, ARC-AGI-2, LiveBench, and AA-AnalystAgent — are independent runs, not Anthropic's own reported figures: HLE, Terminal-Bench 2.1 and AA-AnalystAgent are from Artificial Analysis (observed 2026-08-19, 2026-08-20 and 2026-09-29), DeepSWE is from DataCurve (observed 2026-08-17), Agents' Last Exam is from Snorkel AI (re-read 2026-10-01), SWE-bench Verified, LiveCodeBench, and GPQA Diamond are all from Vals.ai (observed 2026-08-17 and 2026-08-20), and ARC-AGI-2 and LiveBench are from ARC Prize and LiveBench's own board (both observed 2026-08-24).
Does 'half the price' mean Claude Opus 5 is cheaper than Claude Opus 4.8?
No. Claude Opus 5 is priced at $5 per million input tokens and $25 per million output tokens — identical to Claude Opus 4.8's rate, per Anthropic's pricing docs. The 'half the price' language in Anthropic's launch messaging compares Opus 5 to Claude Fable 5, priced at $10/$50 per million tokens on the same Anthropic models-overview page — a separate and pricier model in the current lineup, not a discount off its own Opus predecessor.
What is the effort toggle on Claude Opus 5, and which level is the default?
Claude Opus 5 carries the configurable 'effort' parameter — low, medium, high, xhigh, or max — first introduced with the full five-level scale on Claude Opus 4.7 in April 2026. It trades reasoning depth and token spend against speed and cost. One thing Opus 5 does change: at xhigh or max, thinking can no longer be disabled at all, a breaking change from Opus 4.8. Anthropic's documentation sets the default to high on the Claude API and in Claude Code; that's narrower than Claude Opus 4.8, where high is the default across every surface including claude.ai, so effort may need to be set explicitly elsewhere for Opus 5.
Can Claude Opus 5 generate more than 128,000 output tokens in one response?
Not through the standard Messages API, which caps Claude Opus 5 at 128K output tokens. Anthropic's Message Batches API supports up to 300K output tokens for Opus 5 using the `output-300k-2026-03-24` beta header — the batch-only extended-output mode this model's notes refer to.
Does Claude Opus 5 support the same extended-thinking controls as older Claude models?
No. Claude Opus 5 replaces the older manual 'extended thinking' toggle (with a fixed token budget) with adaptive thinking, tuned instead via the effort parameter. Anthropic's own model comparison table lists extended thinking (`thinking.type: "enabled"`) as unsupported on Opus 5, with adaptive thinking supported in its place.
Further reading
- GPQA Diamond leaderboard 2026 — Claude Opus 5 is one of the 23 models it compares.
- Models with 10M token context windows 2026 — Claude Opus 5 is one of the 35 models it compares.