Anthropic

Claude Opus 5.5

Anthropic's "new leading model" by its own description, at $4/$20 per million tokens — 20% below Opus 5 per token, 40% of Claude Fable 5.1's price. Independently, it edges Fable 5.1 on Humanity's Last Exam (61.4% vs 59.1%, just past the noise band — the highest HLE score Artificial Analysis has recorded) and ties it on LiveBench (83.2 vs 83.4). Nine of the eleven boards hadn't scored it yet.

Claude Opus 5.5 benchmarks and pricing, every number sourced: 2 tracked Claude Opus 5.5 benchmark scores (2 independently run, 0 still resting on a vendor’s own claim), priced at $4.00 per million input tokens and $20.00 per million output.

Claude Opus 5.5’s 2 benchmark scores on this page were verified against their source on 2026-09-23.

Released
2026-09-22
License
proprietary
Context window
1M tokens
Knowledge cutoff
2026-06
Parameters
Not disclosed
Architecture
Not disclosed

Claude Opus 5.5’s verified record

Claude Opus 5.5’s most-compared rival is Claude Opus 5: 2 leads across their 2 shared comparisons. Claude Opus 5.5 is priced at $4.00/$20.00 per 1M tokens (in/out) vs Claude Opus 5’s $5.00/$25.00.

Against the 52 head-to-head comparisons Claude Opus 5.5 shares with other tracked models: 47 real gaps, 4 inside the noise band, and 1 we will not call.

A gap counts for Claude Opus 5.5 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Claude Opus 5.5 trails on 0 of them.

HLE · no tools Reasoning

±2 is noise
Ahead
Claude Fable 5.1 +2.3 · Claude Opus 5 +6.5 · GPT-6 Astra +6.7 · GPT-5.6 Sol +11.9 · MiMo-V2.6-Pro +12.0 · Muse Spark 1.3 +12.3 · GPT-6 Sol +13.5 · Gemini 3.8 Flash +13.6 · Gemini 3.1 Pro Preview +14.4 · Kimi K3 +14.5 · Step 5 Preview +14.9 · Grok 4.7 +18.3 · Qwen3.8-Max +18.4 · Grok 4.6 +18.5 · GLM-5.3 +19.1 · Claude Sonnet 5 +20.1 · DeepSeek V4 Pro (0813) +20.4 · GLM-5.3-Flash +21.5 · GPT-5.6 Luna +21.9 · GPT-6 Luna +22.9 · Qwen3.8-Flash-Next +23.4 · 6 superseded: Claude Fable 5 +5.9 · Claude Opus 4.8 +12.7 · Gemini 3.7 Flash +13.5 · Muse Spark 1.2 +15.9 · GLM-5.2 +20.3 · DeepSeek V4 Flash (0731) +22.8
Unverified
1 model — vendor-reported on one side

LiveBench Composite score across 7 domains

±2.7 is noise
Ahead
Claude Opus 5 +3.1 · GPT-6 Sol +3.9 · Kimi K3 +4.0 · Qwen3.8-Max +4.7 · Grok 4.6 +5.2 · DeepSeek V4 Pro (0813) +5.8 · Grok 4.7 +5.8 · Gemini 3.1 Pro Preview +6.2 · Qwen3.8-Flash-Next +7.0 · GLM-5.3 +7.1 · Claude Sonnet 5 +7.2 · GPT-5.6 Luna +9.6 · GPT-6 Luna +11.2 · GLM-5.3-Flash +11.6 · MiniMax M3 +15.9 · 5 superseded: Gemini 3.7 Flash +4.4 · Muse Spark 1.2 +5.2 · Claude Opus 4.8 +7.0 · DeepSeek V4 Flash (0731) +9.0 · GLM-5.2 +10.0
Tie
4 models within ±2.7

Claude Opus 5.5 API pricing

$4.00 in / $20.00 out per 1M tokens official pricing

What Claude Opus 5.5 costs per job

Claude Opus 5.5 cost for three reference workloads, computed from its list rates
WorkloadTokens in / outCost
One long chat turn100K / 10K$0.600
A codebase review1,000K / 100K$6.00
A day of agent work10,000K / 1,000K$60.00

Computed from Claude Opus 5.5’s list rates above — cache discounts and batch tiers are not applied.

Claude Opus 5.5 is one of 6 Anthropic models tracked on this site, at these official list prices.

Anthropic model pricing, official list rates
ModelIn / 1MOut / 1M
Claude Opus 5.5$4.00$20.00
Claude Fable 5.1$10.00$50.00
Claude Opus 5$5.00$25.00
Claude Sonnet 5$2.00$10.00
Claude Fable 5(superseded)$10.00$50.00
Claude Opus 4.8(superseded)$5.00$25.00

Claude Opus 5.5 benchmark scores

Claude Opus 5.5 benchmark scores, provenance, and source links
BenchmarkScore
HLE(no tools)[1]
Reasoning · ±2 is noise
LiveBench[2]
Composite score across 7 domains · ±2.7 is noise

Who ran these numbers: 2 of 2 independent — artificialanalysis.ai (1), livebench.ai (1).

  1. HLE: AA's own run at max effort (Adaptive Reasoning, Max Effort, Default Fallback) — text-only 2,158-question subset. Currently the highest HLE score AA has recorded for any model.
  2. LiveBench: Board row 'Claude 5.5 Opus Thinking (Max Effort)' on the 2026-06-25 LiveBench release; second on the raw board behind Claude Fable 5.1 (Max Effort, 83.4), a 0.2-point gap inside LiveBench's 2.7-point noise band.

Notes on the record

Pricing is flat — Anthropic's own pricing page states the full 1M-token window bills at the same per-token rate regardless of prompt length, no threshold. $4 input / $20 output per million tokens, cache read at $0.20 (a 0.05x multiplier, lower than Opus 5's 0.1x), cache write $5 (5-minute) or $8 (1-hour), batch API $2/$10, and a research-preview "fast" mode at $8/$40. US-only data residency adds a flat 10% surcharge across all token categories — not a length tier. Anthropic also claims Opus 5.5 costs 40% less than Opus 5 on typical workloads at default settings, combining the lower price with lower token use — its own test, not independently measured.

Context window 1,000,000 tokens, max output 128,000 (extendable to 300,000 via a Message Batches API beta). Input: text and images. Output: text only. Knowledge cutoff June 2026, one month newer than Opus 5's May 2026. License not stated on any Anthropic page checked; inferred proprietary from API-only distribution, consistent with the rest of the lineup.

No page uses "supersede" or "replace" for Opus 5.5's relationship to Opus 5. Opus 5's own model page carries a "Legacy" badge and a migration recommendation — but Anthropic's separate model-deprecations page lists claude-opus-5 as "Active," not "Legacy" or "Deprecated." Two official Anthropic sources disagree on the same model's status; recorded as-is, not resolved.

Anthropic's announcement calls Opus 5.5 "the new leading model" that "performs at the level of Claude Fable 5.1 on most work," while Fable 5.1 (shipped 2026-09-01) stays on sale at $10/$50. This site's independent data agrees on the narrow part it can test: Opus 5.5 leads Fable 5.1 on Humanity's Last Exam by 2.3 points, just outside that benchmark's noise band, and the two tie on LiveBench (83.2 vs 83.4).

Compare with

FAQ

Has Claude Opus 5.5 been independently benchmarked?

Only partly, one day after launch — two of the eleven benchmarks this site tracks carry an independently-run score: Humanity's Last Exam (61.4%, Artificial Analysis, text-only, max effort — the highest AA has recorded for any model) and LiveBench (83.2 overall, max effort, level with Claude Fable 5.1's 83.4 inside the noise band). GPQA Diamond, SWE-bench Verified, LiveCodeBench, Terminal-Bench 2.1, DeepSWE, ARC-AGI-2, Toolathlon-Verified, Agents' Last Exam and HMMT Feb 2026 have none yet — deepswe.datacurve.ai, arcprize.org, tbench.ai, toolathlon.xyz, snorkel.ai, matharena.ai and vals.ai were all checked directly.

Does Claude Opus 5.5 replace Claude Opus 5?

Anthropic's own pages disagree with each other. Opus 5's model page carries a "Legacy" badge and recommends migrating to Opus 5.5, but Anthropic's separate model-deprecations page still lists Opus 5 as "Active," with no deprecation date set. Neither page uses the word "supersede" or "replace." This page records both statements rather than picking one.

Is Claude Opus 5.5 Anthropic's most capable model?

Anthropic doesn't use that phrase, but its announcement calls Opus 5.5 "the new leading model" and says it "performs at the level of Claude Fable 5.1 on most work" — while Fable 5.1 stays on sale at $10/$50 per million tokens, 2.5 times Opus 5.5's $4/$20. On the two benchmarks this site can compare independently, Opus 5.5 leads Fable 5.1 on Humanity's Last Exam by 2.3 points, just past the 2-point noise band, and ties it on LiveBench.

Does Claude Opus 5.5's price change with longer prompts?

No. Anthropic's pricing page states models from Claude 4.6 onward bill the full 1,000,000-token window at one flat rate — a 900,000-token request costs the same per token as a 9,000-token one. The only surcharge is a flat 10% add-on for US-only data residency, which applies regardless of prompt length.

Further reading

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.