Anthropic

Claude Sonnet 5.5

Anthropic's mid-tier Claude 5.5 model at Sonnet 5's unchanged $2/$10 per million tokens; its system card says it "is not at the capability frontier." Two boards score both at the same tier: on Artificial Analysis's Humanity's Last Exam it leads Sonnet 5 by 13.7 points (55.0 vs 41.3), past the 2-point band; on LiveBench, +1.7 (77.8 vs 76.0) is a tie. Nine of twelve boards haven't scored it.

Claude Sonnet 5.5 benchmarks and pricing, every number sourced: 6 tracked Claude Sonnet 5.5 benchmark scores (3 independently run, 3 still resting on a vendor’s own claim), priced at $2.00 per million input tokens and $10.00 per million output.

Claude Sonnet 5.5’s 6 benchmark scores on this page were verified against their source on 2026-09-29.

Released
2026-09-28
License
proprietary
Context window
1M tokens
Knowledge cutoff
2026-06
Parameters
Not disclosed
Architecture
Not disclosed

Claude Sonnet 5.5’s verified record

Claude Sonnet 5.5’s most-compared rival is DeepSeek V4 Pro (0813): 1 lead, 1 tie, and 4 not callable across their 6 shared comparisons. Claude Sonnet 5.5 is priced at $2.00/$10.00 per 1M tokens (in/out) vs DeepSeek V4 Pro (0813)’s $1.32/$3.96.

Against the 135 head-to-head comparisons Claude Sonnet 5.5 shares with other tracked models: 36 real gaps, 17 inside the noise band, and 82 we will not call.

A gap counts for Claude Sonnet 5.5 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Claude Sonnet 5.5 trails on 7 of them.

HLE · no tools Reasoning

±2 is noise
Behind
Ahead
GPT-5.6 Sol +5.5 · MiMo-V2.6-Pro +5.6 · Muse Spark 1.3 +6.3 · GPT-6 Sol +7.1 · Gemini 3.8 Flash +7.2 · Gemini 3.1 Pro Preview +8.0 · Kimi K3 +8.1 · Step 5 Preview +8.5 · Qwen3.8-Max +11.9 · Grok 4.7 +11.9 · Grok 4.6 +12.1 · GLM-5.3 +12.7 · Claude Sonnet 5 +13.7 · DeepSeek V4 Pro (0813) +14.0 · GLM-5.3-Flash +15.1 · GPT-5.6 Luna +15.5 · GPT-6 Luna +16.5 · Qwen3.8-Flash-Next +17.0 · 5 superseded: Claude Opus 4.8 +6.3 · Gemini 3.7 Flash +7.1 · Muse Spark 1.2 +9.5 · GLM-5.2 +13.9 · DeepSeek V4 Flash (0731) +16.4
Tie
3 models within ±2
Unverified
2 models — vendor-reported on one side

LiveBench Composite score across 7 domains

±2.7 is noise
Behind
GPT-5.6 Sol −3.2 · Muse Spark 1.3 −3.8 · Claude Opus 5.5 −5.4 · Claude Fable 5.1 −5.6 · 1 superseded: Claude Fable 5 −5.2
Ahead
GPT-5.6 Luna +4.2 · GPT-6 Luna +5.8 · GLM-5.3-Flash +6.2 · MiniMax M3 +10.5 · 2 superseded: DeepSeek V4 Flash (0731) +3.6 · GLM-5.2 +4.6
Tie
14 models within ±2.7

No verdict for Claude Sonnet 5.5 anywhere on DeepSWE, HLE · with tools, Terminal-Bench 2.1, Toolathlon-Verified (nothing independently confirmed on both sides).

  • None of Claude Sonnet 5.5’s coding comparisons are independently confirmed on both sides yet.
  • None of Claude Sonnet 5.5’s agentic comparisons are independently confirmed on both sides yet.

Claude Sonnet 5.5 API pricing

$2.00 in / $10.00 out per 1M tokens — official pricing

What Claude Sonnet 5.5 costs per job

Claude Sonnet 5.5 cost for three reference workloads, computed from its list rates
WorkloadTokens in / outCost
One long chat turn100K / 10K$0.300
A codebase review1,000K / 100K$3.00
A day of agent work10,000K / 1,000K$30.00

Computed from Claude Sonnet 5.5’s list rates above — cache discounts and batch tiers are not applied.

Claude Sonnet 5.5 is one of 7 Anthropic models tracked on this site, at these official list prices.

Anthropic model pricing, official list rates
ModelIn / 1MOut / 1M
Claude Sonnet 5.5$2.00$10.00
Claude Opus 5.5$4.00$20.00
Claude Fable 5.1$10.00$50.00
Claude Opus 5$5.00$25.00
Claude Sonnet 5$2.00$10.00
Claude Fable 5(superseded)$10.00$50.00
Claude Opus 4.8(superseded)$5.00$25.00

Claude Sonnet 5.5 benchmark scores

Claude Sonnet 5.5 benchmark scores, provenance, and source links
BenchmarkScore
HLE(no tools)[1]
Reasoning · ±2 is noise
LiveBench[2]
Composite score across 7 domains · ±2.7 is noise
Terminal-Bench 2.1[3]
Terminal ops · ±10.6 is noise
HLE(with tools)[4]
Reasoning · ±2 is noise
DeepSWE[5]
Long-horizon coding · ±9.5 is noise
Toolathlon-Verified[6]
Multi-tool chores · ±9.7 is noise

Who ran these numbers: 3 of 6 independent — artificialanalysis.ai (1), livebench.ai (1), vals.ai (1); vendor self-reported (3).

  1. HLE: AA's own run, board row 'Claude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)' — text-only 2,158-question subset, underlying 0.5496. Anthropic's own no-tools figure is 56.9 (system card). AA also lists xhigh 50.0, high 45.8, medium 39.8 and low 36.2 — not tracked here.
  2. LiveBench: LiveBench's default-displayed 'Claude Sonnet 5.5 xHigh Effort' row on the 2026-06-25 release (overall 77.75); a 'Max Effort' variant, 75.7, sits behind the board's variants toggle. Sonnet 5's only row is xHigh Effort (76.0), so the two compare at the same tier. Sub-scores (xHigh): Reasoning 86.8, Coding 88.9, Agentic Coding 39.3, Mathematics 96.7, Data Analysis 78.6, Language 83.4, Instruction Following 70.5.
  3. Terminal-Bench 2.1: vals.ai's own Terminal-Bench 2.1 run (Terminus 2 harness on Daytona), 83.146 ± 1.716, at high effort (vals runs its other benchmarks at max) with Claude Sonnet 5 as server-side fallback for refusals; 80.52 if fallback-assisted tasks are counted as failures. Neither tbench.ai's official board nor Artificial Analysis had run the model as of 2026-09-29, so this is the only independent 2.1 number — comparisons against tbench.ai or AA rows read Setup-dependent.
  4. HLE: Anthropic's launch table at max effort: web search, web fetch, programmatic tool calling and code execution, 980k-token task budget, Claude Opus 4.6 as grader (system card). Sonnet 5 in the same table: 54.9; Opus 5.5: 67.7.
  5. DeepSWE: System card §8.3: average over five trials on the 113-task DeepSWE v1.1 set; the card states no effort level for this run. Not on deepswe.datacurve.ai's board as of 2026-09-29, where Sonnet 5's independent row reads 54.
  6. Toolathlon-Verified: System card §8.14.5: pass@1 77.8 on Anthropic's internal harness at max effort, three trials; two of 324 trials were stopped by production safety classifiers and counted as failures. The same table gives Sonnet 5 74.7 and Opus 5.5 77.8. Not on toolathlon.xyz's board as of 2026-09-29, where Sonnet 5's official row reads 71.6.

Notes on the record

Pricing $2 in / $10 out per million tokens, identical to Sonnet 5 including cache and batch rates (per Anthropic's What's-new page); the full 1,000,000-token window bills at one rate. Knowledge cutoff June 2026; same tokenizer as Sonnet 5. Five effort levels, default high on the API and medium in Claude Code, recalibrated from Sonnet 5's. No parameter count disclosed; license not stated, inferred proprietary.

Anthropic calls it "a faster, lower-cost complement to Claude Opus 5.5" — 30%+ faster output and up to 30% less per task than Sonnet 5, its own testing — and its system card says it "is not at the capability frontier." No lifecycle statement uses "supersede" or "replace" for Sonnet 5: that model's page now carries a "Legacy" badge and a migration recommendation, the models overview lists it under legacy models, and the deprecations page lists it "Active" — the same disagreement recorded for Opus 5.

Against Sonnet 5, two clean independent comparisons exist. Artificial Analysis's Humanity's Last Exam at max effort: 55.0 vs 41.3, a 13.7-point real gap close to Anthropic's own table (56.9 vs 43.1). LiveBench at xHigh: 77.8 vs 76.0, a tie inside the 2.7 band, with Coding, Mathematics, Data Analysis and Instruction Following up and Reasoning and Agentic Coding (39.3 vs 59.4) down. vals.ai's Terminal-Bench 2.1 run reads 83.15 vs 74.53, inside the 10.6 band, with 5.5 at high effort. Anthropic's largest claimed jump is on a benchmark this site doesn't track: Terminal-Bench 4.0, 70.6 vs 10.3.

Provenance caveat: cyber and frontier-LLM safety blocks fall back visibly to Sonnet 5 (apps by default, API by opt-in) — 1.2% of requests in Anthropic's own Terminal-Bench 4.0 run, 2.27% in vals.ai's testing, and Artificial Analysis labels its rows "Default Fallback" — so a published score can contain Sonnet 5 answers. The three self-reported rows here are system-card figures (DeepSWE's effort level unstated); Anthropic's own no-tools HLE, 56.9, sits 1.9 above AA's 55.0.

Compare with

FAQ

Has Claude Sonnet 5.5 been independently benchmarked?

Partly, one day after launch — three of the twelve benchmarks this site tracks carry an independently-run score: Humanity's Last Exam (55.0, Artificial Analysis, max effort, text-only), LiveBench (77.8, the board's default xHigh Effort row; a Max Effort variant reads 75.7) and Terminal-Bench 2.1 (83.15, vals.ai's own run at high effort — neither tbench.ai's board nor Artificial Analysis has run it). GPQA Diamond, SWE-bench Verified and LiveCodeBench will not get rows from vals.ai, which has archived those boards as saturated, and Artificial Analysis has published no GPQA Diamond or Terminal-Bench 2.1 score for any model released after 2026-09-11. DeepSWE, Toolathlon-Verified, Agents' Last Exam, ARC-AGI-2, HMMT Feb 2026 and AA-AnalystAgent had no row as of 2026-09-29.

Is Claude Sonnet 5.5 better than Claude Sonnet 5?

On the two same-tier independent comparisons, one real gap and one tie. Humanity's Last Exam: 55.0 vs 41.3 (Artificial Analysis, both at max effort), a 13.7-point gap against a 2-point noise band. LiveBench: 77.8 vs 76.0 (both xHigh Effort), a 1.7-point gap inside the 2.7 band, with Coding 88.9 vs 80.7, Mathematics 96.7 vs 92.9, Data Analysis 78.6 vs 71.7 and Instruction Following 70.5 vs 63.9 up, Reasoning 86.8 vs 88.7 and Agentic Coding 39.3 vs 59.4 down. Anthropic's own table shows its largest gains on benchmarks this site doesn't track, such as Terminal-Bench 4.0 (70.6 vs 10.3). The price is unchanged, so the per-task saving Anthropic claims, up to 30%, rests on its own speed and token-use measurements.

Does Claude Sonnet 5.5 replace Claude Sonnet 5?

Anthropic's pages disagree, exactly as they did for Opus 5. Sonnet 5's model page carries a "Legacy" badge and says to consider migrating to Sonnet 5.5, the models overview lists Sonnet 5 under legacy models, and the deprecations page lists it "Active" with no retirement before 2027-06-30. No lifecycle statement uses "supersede" or "replace" (the migration guide's "replace your model ID" is a code instruction), so this site keeps Sonnet 5 current and records all three statements. Sonnet 5 also has a new job: it is the visible fallback model when Sonnet 5.5's cyber or frontier-LLM safety classifiers block a request.

How does Claude Sonnet 5.5 compare with Claude Opus 5.5?

Anthropic positions it as a complement, not a rival: "not at the capability frontier" and "less capable than Claude Opus 5.5 on most evaluations we report," per the system card. The two independent boards that have scored both at max effort agree: Opus 5.5 leads Humanity's Last Exam 61.4 vs 55.0, a 6.4-point real gap, and LiveBench 83.2 vs 77.8, a 5.4-point real gap (Opus 5.5's row is Max Effort; its xHigh row, 82.1, still clears the band) — at twice the price, $4/$20 against $2/$10 per million tokens.

Does Claude Sonnet 5.5's price change with longer prompts?

No. Sonnet 5.5 inherits the rule Anthropic applies to every model from Claude 4.6 onward: the whole 1,000,000-token window is billed at the standard $2/$10 per million tokens, with no threshold above which the rate rises. The only multipliers are 1.1x for US-only inference and the platform premiums Amazon Bedrock and Google Cloud charge on regional endpoints — neither depends on prompt length.

Can a Claude Sonnet 5.5 benchmark score include answers from Claude Sonnet 5?

Yes, on some setups. Anthropic's system card says cyber and frontier-LLM safety blocks fall back to Sonnet 5 in its first-party products and for API developers who opt in, and it reports 1.2% of requests in its own Terminal-Bench 4.0 run were answered by the fallback. vals.ai states it ran Sonnet 5.5 with Sonnet 5 as a server-side fallback (a 2.27% fallback rate) and publishes a second figure with fallback-assisted tasks counted as failures; Artificial Analysis labels its rows "Default Fallback" without publishing a rate. Biology, reasoning-extraction and general-harm blocks have no fallback and simply fail.

Further reading

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.