Anthropic

Claude Opus 5

Anthropic pitches Opus 5 as "close to the frontier intelligence of Claude Fable 5 at half the price" — but that price is just Claude Opus 4.8's unchanged $5/$25 rate carried over, not a cut. What's real is the scorecard: all ten benchmark results tracked here are independent runs, not Anthropic's own launch-day figures.

Claude Opus 5 benchmarks and pricing, every number sourced: 10 tracked Claude Opus 5 benchmark scores (10 independently run, 0 still resting on a vendor’s own claim), priced at $5.00 per million input tokens and $25.00 per million output.

Claude Opus 5’s 10 benchmark scores on this page were each verified against their sources between 2026-08-17 and 2026-10-01.

Version history: succeeded Claude Opus 4.8 (2026-05-28).

Released
2026-07-24
License
proprietary
Context window
1M tokens
Knowledge cutoff
2026-05
Parameters
Not disclosed
Architecture
Not disclosed

Claude Opus 5’s verified record

Claude Opus 5’s featured comparison is Claude Sonnet 5: 3 leads, 1 tie, and 4 not callable. Claude Opus 5 is priced at $5.00/$25.00 per 1M tokens (in/out) vs Claude Sonnet 5’s $2.00/$10.00. Full verdict →

Against the 214 head-to-head comparisons Claude Opus 5 shares with other tracked models: 65 real gaps, 53 inside the noise band, and 96 we will not call.

A gap counts for Claude Opus 5 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Claude Opus 5 trails on 7 of them.

HLE · no tools Reasoning

±2 is noise
Behind
Ahead
GPT-6.1 Sol +2.0 · GPT-5.6 Sol +5.4 · MiMo-V2.6-Pro +5.5 · Muse Spark 1.3 +6.2 · GPT-6 Sol +7.0 · Gemini 3.8 Flash +7.1 · Gemini 3.1 Pro Preview +7.9 · Kimi K3 +8.0 · Step 5 Preview +8.4 · Qwen3.8-Max +11.8 · Grok 4.7 +11.8 · Grok 4.6 +12.0 · GLM-5.3 +12.6 · Claude Sonnet 5 +13.6 · DeepSeek V4 Pro (0813) +13.9 · GLM-5.3-Flash +15.0 · GPT-5.6 Luna +15.4 · GPT-6 Luna +16.4 · Qwen3.8-Flash-Next +16.9 · 5 superseded: Claude Opus 4.8 +6.2 · Gemini 3.7 Flash +7.0 · Muse Spark 1.2 +9.4 · GLM-5.2 +13.8 · DeepSeek V4 Flash (0731) +16.3
Tie
3 models within ±2
Unverified
2 models — vendor-reported on one side

LiveBench Composite score across 7 domains

±2.7 is noise
Behind
Claude Opus 5.5 −3.1 · Claude Fable 5.1 −3.3 · 1 superseded: Claude Fable 5 −2.9
Ahead
Gemini 3.1 Pro Preview +3.1 · Qwen3.8-Flash-Next +3.9 · GLM-5.3 +4.0 · Claude Sonnet 5 +4.1 · GPT-5.6 Luna +6.5 · GPT-6 Luna +8.1 · GLM-5.3-Flash +8.5 · MiniMax M3 +12.8 · 3 superseded: Claude Opus 4.8 +3.9 · DeepSeek V4 Flash (0731) +5.9 · GLM-5.2 +6.9
Tie
12 models within ±2.7

Agents' Last Exam Professional work

±3.2 is noise
Behind
Ahead
Kimi K3 +3.9 · Qwen3.8-Max +5.2 · GPT-6 Luna +7.2 · Gemini 3.1 Pro Preview +15.8 · 3 superseded: Claude Opus 4.8 +5.2 · Claude Fable 5 +6.5 · GLM-5.2 +11.8
Tie
5 models within ±3.2
Unverified
6 models — vendor-reported on one side

DeepSWE Long-horizon coding

±9.5 is noise
Ahead
GLM-5.3-Flash +11.0 · Qwen3.8-Max +17.0 · Claude Sonnet 5 +20.0 · 4 superseded: Claude Opus 4.8 +15.0 · Muse Spark 1.2 +19.0 · DeepSeek V4 Flash (0731) +21.0 · GLM-5.2 +30.0
Tie
9 models within ±9.5
Unverified
10 models — vendor-reported on one side
Setup-dependent
2 models — scored on a different harness

ARC-AGI-2 · max Compositional visual reasoning

±9.2 is noise
Ahead
DeepSeek V4 Pro (0813) +29.1 · Kimi K3 +30.0 · GPT-5.6 Luna +30.8 · GPT-6 Luna +31.1 · 1 superseded: DeepSeek V4 Flash (0731) +29.0
Tie
4 models within ±9.2

AnalystAgent Spreadsheet & document analysis

±11.2 is noise
Ahead
Gemini 3.1 Pro Preview +12.5 · Grok 4.6 +12.5 · Kimi K3 +15.0 · MiniMax M3 +43.8
Tie
7 models within ±11.2

No verdict for Claude Opus 5 anywhere on Terminal-Bench 2.1 (every independently confirmed comparison inside the noise band); GPQA Diamond, LiveCodeBench, SWE-bench Verified (saturated).

What changed from Claude Opus 4.8 to Claude Opus 5

The 9 benchmarks both models have been scored on, using the same variant each time. A raw Claude Opus 5 gain is not a real gain until it clears that benchmark’s own noise band, so each row below carries the verdict and not just the arithmetic.

Claude Opus 5 versus Claude Opus 4.8, per-benchmark change and whether it clears the noise band
BenchmarkClaude Opus 4.8Claude Opus 5ChangeVerdict
HLE48.754.9+6.2Real gap
Terminal-Bench 2.178.989.1+10.2Setup-dependent
DeepSWE5974+15.0Real gap
Agents' Last Exam2732.2+5.2Real gap
SWE-bench Verified88.697+8.4Tainted
GPQA Diamond92.4293.43+1.0Tainted
LiveCodeBench87.8289.03+1.2Tainted
LiveBench76.280.1+3.9Real gap
AnalystAgent4553.75+8.8Tie

Claude Opus 5 API pricing

$5.00 in / $25.00 out per 1M tokens — official pricing

What Claude Opus 5 costs per job

Claude Opus 5 cost for three reference workloads, computed from its list rates
WorkloadTokens in / outCost
One long chat turn100K / 10K$0.750
A codebase review1,000K / 100K$7.50
A day of agent work10,000K / 1,000K$75.00

Computed from Claude Opus 5’s list rates above — cache discounts and batch tiers are not applied.

Claude Opus 5 is one of 7 Anthropic models tracked on this site, at these official list prices.

Anthropic model pricing, official list rates
ModelIn / 1MOut / 1M
Claude Sonnet 5.5$2.00$10.00
Claude Opus 5.5$4.00$20.00
Claude Fable 5.1$10.00$50.00
Claude Opus 5$5.00$25.00
Claude Sonnet 5$2.00$10.00
Claude Fable 5(superseded)$10.00$50.00
Claude Opus 4.8(superseded)$5.00$25.00

Claude Opus 5 benchmark scores

Claude Opus 5 benchmark scores, provenance, and source links
BenchmarkScore
HLE(no tools)[1]
Reasoning · ±2 is noise
Terminal-Bench 2.1[2]
Terminal ops · ±10.6 is noise
DeepSWE[3]
Long-horizon coding · ±9.5 is noise
SWE-bench Verifiedsaturated[4]
Bug fixing — not ranked at any gap size
LiveCodeBenchsaturated[5]
Contest coding — not ranked at any gap size
GPQA Diamondsaturated[6]
Expert science Q&A — not ranked at any gap size
Agents' Last Exam[7]
Professional work · ±3.2 is noise
ARC-AGI-2(max)[8]
Compositional visual reasoning · ±9.2 is noise
LiveBench[9]
Composite score across 7 domains · ±2.7 is noise
AnalystAgent[10]
Spreadsheet & document analysis · ±11.2 is noise

Who ran these numbers: 10 of 10 independent — artificialanalysis.ai (3), vals.ai (3), deepswe.datacurve.ai (1), snorkel.ai (1), arcprize.org (1), livebench.ai (1).

  1. HLE: AA live leaderboard value, 'Claude Opus 5 (Adaptive Reasoning, Max Effort)' — text-only, no tools. Re-verified on the board 2026-08-19. Corrected 2026-08-19: this row used to cite AA's Opus 5 launch article, which reports 53.0 — the 54.9 was always the board's number, so the citation pointed at a page that did not support it.
  2. Terminal-Bench 2.1: AA's own run, chart-JSON label 'Claude Opus 5 (max)'; page prose gives the fuller label 'Claude Opus 5 (Adaptive Reasoning, Max Effort)' for the same score. A lower-effort 'Claude Opus 5 (xhigh)' variant also appears in the same chart at 88.0. Not on tbench.ai's official board, which lists only older Anthropic models (Fable 5, Opus 4.8, Opus 4.7, Sonnet 5). Cross-check (2026-10-01): vals.ai's archived Terminus 2 table lists it at 84.64, with Opus 4.8 as a refusal fallback on 9 of 267 tasks; counting those as failures gives 81.27.
  3. DeepSWE: 74%±4 pass rate on the mini-swe-agent harness (all models run identically); board updated 2026-08-13.
  4. SWE-bench Verified: vals.ai run, rank 1 of 83 (updated 2026-08-14); Anthropic self-reports 96.0. Benchmark near-saturated — five models at 95%+.
  5. LiveCodeBench: vals.ai, updated 2026-08-15; re-verified on the vals.ai model page 2026-08-17.
  6. GPQA Diamond: vals.ai run, rank 7/135 (board updated 2026-08-19).
  7. Agents' Last Exam: Overall pass rate 32.2 (Claude Code, High, the board's best effort for this model; score 55.9; est. cost $1,108). Per-effort results on the board: High 32.2, Max 30.9, XHigh 30.3, Medium 28.9, Low 27.6. Snorkel's data notes record that the board's upstream source re-published every Claude Opus 5 figure on 2026-09-24, moving its Overall best tier (High) from 31.6 to 32.2. This site's earlier 27.0 (Max; score 49.0; est. cost $7,239, recorded 2026-08-20) no longer appears; the current Max run reads 30.9 (score 52.7).
  8. ARC-AGI-2: Claude Opus 5's official ARC-AGI-2 leaderboard row, dated 2026-07-24 on arcprize.org. Higher than its only other populated tier, High (88.3%).
  9. LiveBench: Board row "Claude 5 Opus Thinking Max Effort (LiveBench's own row label reverses this site's word order; corroborated as the same model by a real GitHub issue on livebench/livebench)" on the 2026-06-25 LiveBench release.
  10. AnalystAgent: AA's own run, board row 'Claude Opus 5 (Adaptive Reasoning, Max Effort)'; the board rounds this to 53.8. pass@1 63.75, pass@5 73.75.

Notes on the record

Successor to Claude Opus 4.8 at the same $5/$25 price (confirmed current on Anthropic's pricing/model overview page as of 2026-08-20). Configurable 'effort' toggle (shared machinery, first shipped on Claude Opus 4.7) — Anthropic's docs list five levels (low, medium, high, xhigh, max), defaulting to high on the Claude API and Claude Code (source: platform.claude.com/docs/en/about-claude/models/overview). Anthropic positions it as 'close to the frontier intelligence of Claude Fable 5 at half the price'.

Max output 128k (300k via batch beta, using the `output-300k-2026-03-24` beta header on the Message Batches API, per the same source). Also available via Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry.

Compare with

FAQ

Has Claude Opus 5 been independently benchmarked?

Yes, on every score this site currently tracks for it. All ten results — Humanity's Last Exam (no tools), Terminal-Bench 2.1, DeepSWE, SWE-bench Verified, LiveCodeBench, GPQA Diamond, Agents' Last Exam, ARC-AGI-2, LiveBench, and AA-AnalystAgent — are independent runs, not Anthropic's own reported figures: HLE, Terminal-Bench 2.1 and AA-AnalystAgent are from Artificial Analysis (observed 2026-08-19, 2026-08-20 and 2026-09-29), DeepSWE is from DataCurve (observed 2026-08-17), Agents' Last Exam is from Snorkel AI (re-read 2026-10-01), SWE-bench Verified, LiveCodeBench, and GPQA Diamond are all from Vals.ai (observed 2026-08-17 and 2026-08-20), and ARC-AGI-2 and LiveBench are from ARC Prize and LiveBench's own board (both observed 2026-08-24).

Does 'half the price' mean Claude Opus 5 is cheaper than Claude Opus 4.8?

No. Claude Opus 5 is priced at $5 per million input tokens and $25 per million output tokens — identical to Claude Opus 4.8's rate, per Anthropic's pricing docs. The 'half the price' language in Anthropic's launch messaging compares Opus 5 to Claude Fable 5, priced at $10/$50 per million tokens on the same Anthropic models-overview page — a separate and pricier model in the current lineup, not a discount off its own Opus predecessor.

What is the effort toggle on Claude Opus 5, and which level is the default?

Claude Opus 5 carries the configurable 'effort' parameter — low, medium, high, xhigh, or max — first introduced with the full five-level scale on Claude Opus 4.7 in April 2026. It trades reasoning depth and token spend against speed and cost. One thing Opus 5 does change: at xhigh or max, thinking can no longer be disabled at all, a breaking change from Opus 4.8. Anthropic's documentation sets the default to high on the Claude API and in Claude Code; that's narrower than Claude Opus 4.8, where high is the default across every surface including claude.ai, so effort may need to be set explicitly elsewhere for Opus 5.

Can Claude Opus 5 generate more than 128,000 output tokens in one response?

Not through the standard Messages API, which caps Claude Opus 5 at 128K output tokens. Anthropic's Message Batches API supports up to 300K output tokens for Opus 5 using the `output-300k-2026-03-24` beta header — the batch-only extended-output mode this model's notes refer to.

Does Claude Opus 5 support the same extended-thinking controls as older Claude models?

No. Claude Opus 5 replaces the older manual 'extended thinking' toggle (with a fixed token budget) with adaptive thinking, tuned instead via the effort parameter. Anthropic's own model comparison table lists extended thinking (`thinking.type: "enabled"`) as unsupported on Opus 5, with adaptive thinking supported in its place.

Further reading

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.