Anthropic

Claude Sonnet 5

Sonnet 5's $2/$10 launch pricing was introductory, scheduled to rise to $3/$15 on 2026-09-01 — Anthropic cancelled that increase, so $2/$10 is now the permanent rate. All nine benchmark results on this page are independent third-party runs, not Anthropic self-reports — but Sonnet 5's newer tokenizer produces roughly 30% more tokens for the same text, so the effective cost-per-task falls by less than the sticker price alone suggests.

Claude Sonnet 5 benchmarks and pricing, every number sourced: 9 tracked Claude Sonnet 5 benchmark scores (9 independently run, 0 still resting on a vendor’s own claim), priced at $2.00 per million input tokens and $10.00 per million output.

Claude Sonnet 5’s 9 benchmark scores on this page were each verified against their sources between 2026-08-17 and 2026-09-29.

Released
2026-06-30
License
proprietary
Context window
1M tokens
Knowledge cutoff
2026-01
Parameters
Not disclosed
Architecture
Not disclosed

Claude Sonnet 5’s verified record

Claude Sonnet 5’s featured comparison is Claude Opus 5: 3 trails, 1 tie, and 4 not callable. Claude Sonnet 5 is priced at $2.00/$10.00 per 1M tokens (in/out) vs Claude Opus 5’s $5.00/$25.00. Full verdict →

Against the 202 head-to-head comparisons Claude Sonnet 5 shares with other tracked models: 53 real gaps, 40 inside the noise band, and 109 we will not call.

A gap counts for Claude Sonnet 5 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Claude Sonnet 5 trails on 42 of them.

HLE · no tools Reasoning

±2 is noise
Behind
Step 5 Preview −5.2 · Kimi K3 −5.6 · Gemini 3.1 Pro Preview −5.7 · Gemini 3.8 Flash −6.5 · GPT-6 Sol −6.6 · Muse Spark 1.3 −7.4 · MiMo-V2.6-Pro −8.1 · GPT-5.6 Sol −8.2 · GPT-6.1 Sol −11.6 · GPT-6 Astra −13.4 · Claude Opus 5 −13.6 · Claude Sonnet 5.5 −13.7 · Gemini 4 Argon −15.8 · Claude Fable 5.1 −17.8 · Claude Opus 5.5 −20.1 · 4 superseded: Muse Spark 1.2 −4.2 · Gemini 3.7 Flash −6.6 · Claude Opus 4.8 −7.4 · Claude Fable 5 −14.2
Ahead
GPT-6 Luna +2.8 · Qwen3.8-Flash-Next +3.3 · 1 superseded: DeepSeek V4 Flash (0731) +2.7
Tie
8 models within ±2
Unverified
2 models — vendor-reported on one side

LiveBench Composite score across 7 domains

±2.7 is noise
Behind
Kimi K3 −3.2 · GPT-6 Sol −3.3 · Claude Opus 5 −4.1 · GPT-5.6 Sol −5.0 · Muse Spark 1.3 −5.6 · GPT-6.1 Sol −5.6 · Claude Opus 5.5 −7.2 · Claude Fable 5.1 −7.4 · 2 superseded: Gemini 3.7 Flash −2.8 · Claude Fable 5 −7.0
Ahead
GPT-6 Luna +4.0 · GLM-5.3-Flash +4.4 · MiniMax M3 +8.7 · 1 superseded: GLM-5.2 +2.8
Tie
12 models within ±2.7

DeepSWE Long-horizon coding

±9.5 is noise
Behind
Grok 4.6 −13.0 · GPT-5.6 Luna −13.0 · Kimi K3 −15.0 · GLM-5.3 −15.0 · GPT-5.6 Sol −19.0 · Claude Opus 5 −20.0 · Gemini 3.8 Flash −20.0 · GPT-6 Astra −20.0 · 2 superseded: Gemini 3.7 Flash −11.0 · Claude Fable 5 −16.0
Ahead
1 superseded: GLM-5.2 +10.0
Tie
5 models within ±9.5
Unverified
10 models — vendor-reported on one side
Setup-dependent
2 models — scored on a different harness

AnalystAgent Spreadsheet & document analysis

±11.2 is noise
Behind
Claude Fable 5.1 −11.3 · 1 superseded: Gemini 3.7 Flash −13.8
Ahead
Tie
8 models within ±11.2

Toolathlon-Verified Multi-tool chores

±9.7 is noise
Ahead
Gemini 3.1 Pro Preview +10.5 · 1 superseded: GLM-5.2 +11.7
Tie
4 models within ±9.7
Unverified
10 models — vendor-reported on one side

Terminal-Bench 2.1 Terminal ops

±10.6 is noise
Behind
GPT-6 Astra −12.8
Tie
3 models within ±10.6
Unverified
1 model — vendor-reported on one side
Setup-dependent
26 models — scored on a different harness

No verdict for Claude Sonnet 5 anywhere on GPQA Diamond, LiveCodeBench, SWE-bench Verified (saturated).

Claude Sonnet 5 API pricing

$2.00 in / $10.00 out per 1M tokens — official pricing

What Claude Sonnet 5 costs per job

Claude Sonnet 5 cost for three reference workloads, computed from its list rates
WorkloadTokens in / outCost
One long chat turn100K / 10K$0.300
A codebase review1,000K / 100K$3.00
A day of agent work10,000K / 1,000K$30.00

Computed from Claude Sonnet 5’s list rates above — cache discounts and batch tiers are not applied.

Claude Sonnet 5 is one of 7 Anthropic models tracked on this site, at these official list prices.

Anthropic model pricing, official list rates
ModelIn / 1MOut / 1M
Claude Sonnet 5.5$2.00$10.00
Claude Opus 5.5$4.00$20.00
Claude Fable 5.1$10.00$50.00
Claude Opus 5$5.00$25.00
Claude Sonnet 5$2.00$10.00
Claude Fable 5(superseded)$10.00$50.00
Claude Opus 4.8(superseded)$5.00$25.00

Claude Sonnet 5 benchmark scores

Claude Sonnet 5 benchmark scores, provenance, and source links
BenchmarkScore
HLE(no tools)[1]
Reasoning · ±2 is noise
Terminal-Bench 2.1[2]
Terminal ops · ±10.6 is noise
DeepSWE[3]
Long-horizon coding · ±9.5 is noise
SWE-bench Verifiedsaturated[4]
Bug fixing — not ranked at any gap size
GPQA Diamondsaturated[5]
Expert science Q&A — not ranked at any gap size
LiveCodeBenchsaturated[6]
Contest coding — not ranked at any gap size
Toolathlon-Verified[7]
Multi-tool chores · ±9.7 is noise
LiveBench[8]
Composite score across 7 domains · ±2.7 is noise
AnalystAgent[9]
Spreadsheet & document analysis · ±11.2 is noise

Who ran these numbers: 9 of 9 independent — vals.ai (3), artificialanalysis.ai (2), tbench.ai (1), deepswe.datacurve.ai (1), toolathlon.xyz (1), livebench.ai (1).

  1. HLE: Artificial Analysis's own run (max effort, text-only, no tools; 2,158 questions, pass@1). Not comparable to vendors' with-tools HLE claims.
  2. Terminal-Bench 2.1: Sonnet 5's Terminal-Bench 2.1 row was re-verified 2026-09-03 in a browser after npm run check-sources flagged it: tbench.ai rebuilt the site around 2026-08-27 (new default board Terminal-Bench 4.0, version now a client-side selection on a single URL), so the old deep link no longer serves 2.1 in static HTML — see the claude-opus-4-8 row on this same benchmark for the full mechanism. Value confirmed unchanged: 74.6%, rank 10 of 17, 'Sonnet 5 (high)' via Claude Code (CI now shown as ±3.2%, not ±1.6 — recomputed, not a re-run). Cross-corroborated by vals.ai's independent run at 74.53%.
  3. DeepSWE: 54%±4 pass@1 on the mini-swe-agent harness, rank 13/17; $26.40 avg cost/task. Board updated 2026-08-13.
  4. SWE-bench Verified: 79.60%±1.80, rank 22/83 (updated 2026-08-14). vals.ai's earlier Jun-30 setup gave 75.49 — harness changed since, numbers not directly comparable across dates.
  5. GPQA Diamond: 88.89%±2.22, rank 30/133 (updated 2026-08-15). vals.ai notes the benchmark is largely saturated — 24 models at 90%+.
  6. LiveCodeBench: 82.43%±1.09 on v6, rank 50/138 (updated 2026-08-15).
  7. Toolathlon-Verified: Pass@1 71.6±1.2 on the current Toolathlon-Verified table (Pass@3 83.3, Pass^3 53.7 also shown but not tracked here).
  8. LiveBench: Board row "Claude Sonnet 5 xHigh Effort" on the 2026-06-25 LiveBench release.
  9. AnalystAgent: AA's own run, board row 'Claude Sonnet 5 (Adaptive Reasoning, Max Effort)'; the board rounds to one decimal. pass@1 61.5, pass@5 77.5. Source re-pointed 2026-09-30: AA's AA-AnalystAgent leaderboard page no longer embeds this model in its default selection; its model page still shows 46.25%.

Notes on the record

Mid-tier of the Claude 5 family; Anthropic's docs place its predecessor Sonnet 4.6 in 'Legacy models' (platform.claude.com/docs/en/about-claude/models/overview). Since Claude Sonnet 5.5 shipped on 2026-09-28, Sonnet 5's own model page carries the same "Legacy" badge and a migration recommendation, while Anthropic's model-deprecations page still lists it "Active" with no retirement date before 2027-06-30; it also now serves as the visible fallback model for Sonnet 5.5's cyber and frontier-LLM safety blocks (checked 2026-09-29). Launch's introductory $2/$10 pricing, originally scheduled to rise to $3/$15 on 2026-09-01, was instead made the permanent standard price on 2026-08-10 — the increase will not occur, per the official pricing docs (platform.claude.com/docs/en/about-claude/pricing) and Anthropic's own announcement (x.com/claudeai/status/2086891169217122586).

Adaptive thinking with an effort control (low/medium/high/xhigh/max) that defaults to high; unlike Sonnet 4.6, the old manual extended-thinking mode (thinking: {type: "enabled", budget_tokens: N}) and non-default sampling parameters now return a 400 error — thinking can still be disabled outright via thinking: {type: "disabled"} — and Priority Tier — available on Sonnet 4.6 — is not offered on Sonnet 5 (per Anthropic's "What's new in Claude Sonnet 5" docs).

Max output 128k (up to 300k via the Batch API extended-output beta). The full 1M-token context window bills at standard per-token rates with no long-context surcharge or threshold tier; the newer tokenizer used from Sonnet 5 onward produces roughly 30% more tokens than Sonnet 4.6's for the same text (confirmed on Anthropic's "What's new in Claude Sonnet 5" docs, platform.claude.com/docs/en/about-claude/models/whats-new-sonnet-5#new-tokenizer), so per-task cost savings are smaller in practice than the per-token price cut alone suggests.

Compare with

FAQ

Has Claude Sonnet 5 been independently benchmarked?

Yes — every score tracked for Claude Sonnet 5 on this page (HLE, Terminal-Bench 2.1, DeepSWE, SWE-bench Verified, GPQA Diamond, LiveCodeBench, Toolathlon-Verified, LiveBench, and AA-AnalystAgent) is an independent third-party run — from ArtificialAnalysis, tbench.ai, the DeepSWE board, Vals AI, Toolathlon's own board, or LiveBench's own board — observed 2026-08-17 through 2026-09-29, rather than a number self-reported by Anthropic. That's 9 of 9 listed benchmarks independently run, zero self-reported.

Will Claude Sonnet 5's $2/$10 pricing go up later?

No. The $2-per-million-input / $10-per-million-output rate launched as introductory pricing with a scheduled rise to $3/$15 set for 2026-09-01, but Anthropic canceled that increase on 2026-08-10 and made the launch rate the permanent standard price, per the official pricing docs.

What does the effort setting do on Claude Sonnet 5, and can its thinking be turned off?

Effort (low, medium, high, xhigh, or max) controls how many tokens Claude Sonnet 5 spends reasoning before it answers, and it defaults to high. Thinking itself can still be turned off by passing thinking: {type: "disabled"} — what Anthropic actually removed is the old manual extended-thinking mode (thinking: {type: "enabled", budget_tokens: N}), which now returns a 400 error, per Anthropic's docs.

Is Claude Sonnet 5 a drop-in replacement for Claude Sonnet 4.6?

Mostly. Per Anthropic's docs, Claude Sonnet 5 has three behavior changes from Sonnet 4.6: adaptive thinking now runs by default (it can still be turned off via thinking: {type: "disabled"}), the old manual extended-thinking mode is removed and returns a 400 error, and non-default sampling parameters (temperature, top_p, top_k) also now return a 400 error instead of being silently accepted. Separately, Priority Tier — available on Sonnet 4.6 — is not offered on Sonnet 5.

Is Claude Sonnet 5 cheaper than Claude Opus 5?

Per Anthropic's pricing docs, yes: Sonnet 5 lists at $2/$10 per million input/output tokens versus Opus 5's $5/$25, less than half the per-token price. Because Sonnet 5's tokenizer produces roughly 30% more tokens than Sonnet 4.6's for the same text, the real cost gap on a given task will run narrower than the sticker prices alone suggest.

How much can prompt caching or batch processing cut Claude Sonnet 5's API costs?

Cache reads cost 0.1x the base input price ($0.20 per million tokens), and the Batch API cuts both input and output by 50% ($1/$5 per million tokens), so a cached, batched workload can run well below the $2/$10 sticker price. Routing to US-only inference instead adds a 1.1x multiplier, per Anthropic's pricing docs.

Further reading

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.