AnthropicSuperseded by Claude Opus 5
Claude Opus 4.8
Anthropic's own docs already file Claude Opus 4.8 under "Legacy models" with a migration guide pointing to Opus 5 — but the demotion is mostly cosmetic: it's still formally Active until at least 2027-05-28, priced identically to its successor, and remains the model Anthropic's own API docs use as the worked fallback example when Claude Fable 5's safety classifiers refuse a request.
Claude Opus 4.8 benchmarks and pricing, every number sourced: 13 tracked Claude Opus 4.8 benchmark scores (12 independently run, 1 still resting on a vendor’s own claim), priced at $5.00 per million input tokens and $25.00 per million output.
Claude Opus 4.8’s 13 benchmark scores on this page were each verified against their sources between 2026-08-13 and 2026-09-29.
- Released
- 2026-05-28
- License
- proprietary
- Context window
- 1M tokens
- Knowledge cutoff
- Not disclosed
- Parameters
- Not disclosed
- Architecture
- Not disclosed
Claude Opus 4.8’s verified record
Claude Opus 4.8’s featured comparison is Gemini 3.1 Pro Preview: 3 leads, 3 ties, and 5 not callable. Claude Opus 4.8 is priced at $5.00/$25.00 per 1M tokens (in/out) vs Gemini 3.1 Pro Preview’s $2.00/$12.00. Full verdict →
Against the 234 head-to-head comparisons Claude Opus 4.8 shares with other tracked models: 58 real gaps, 48 inside the noise band, and 128 we will not call.
A gap counts for Claude Opus 4.8 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Claude Opus 4.8 trails on 33 of them.
HLE · no tools Reasoning
±2 is noiseLiveBench Composite score across 7 domains
±2.7 is noiseAgents' Last Exam Professional work
±3.2 is noiseDeepSWE Long-horizon coding
±9.5 is noiseAnalystAgent Spreadsheet & document analysis
±11.2 is noiseToolathlon-Verified Multi-tool chores
±9.7 is noiseTerminal-Bench 2.1 Terminal ops
±10.6 is noiseARC-AGI-2 · high Compositional visual reasoning
±9.2 is noiseNo verdict for Claude Opus 4.8 anywhere on GPQA Diamond, LiveCodeBench, SWE-bench Verified (saturated); HLE · with tools (nothing independently confirmed on both sides); HMMT (contaminated).
Claude Opus 4.8 API pricing
$5.00 in / $25.00 out per 1M tokens — official pricing
What Claude Opus 4.8 costs per job
| Workload | Tokens in / out | Cost |
|---|---|---|
| One long chat turn | 100K / 10K | $0.750 |
| A codebase review | 1,000K / 100K | $7.50 |
| A day of agent work | 10,000K / 1,000K | $75.00 |
Computed from Claude Opus 4.8’s list rates above — cache discounts and batch tiers are not applied.
Claude Opus 4.8 is one of 7 Anthropic models tracked on this site, at these official list prices.
| Model | In / 1M | Out / 1M |
|---|---|---|
| Claude Sonnet 5.5 | $2.00 | $10.00 |
| Claude Opus 5.5 | $4.00 | $20.00 |
| Claude Fable 5.1 | $10.00 | $50.00 |
| Claude Opus 5 | $5.00 | $25.00 |
| Claude Sonnet 5 | $2.00 | $10.00 |
| Claude Fable 5(superseded) | $10.00 | $50.00 |
| Claude Opus 4.8(superseded) | $5.00 | $25.00 |
Claude Opus 4.8 benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise | |
| HLE(with tools)[2] Reasoning · ±2 is noise | |
| Terminal-Bench 2.1[3] Terminal ops · ±10.6 is noise | |
| DeepSWE[4] Long-horizon coding · ±9.5 is noise | |
| Toolathlon-Verified[5] Multi-tool chores · ±9.7 is noise | |
| Agents' Last Exam[6] Professional work · ±3.2 is noise | |
| SWE-bench Verifiedsaturated[7] Bug fixing — not ranked at any gap size | |
| GPQA Diamondsaturated[8] Expert science Q&A — not ranked at any gap size | |
| LiveCodeBenchsaturated[9] Contest coding — not ranked at any gap size | |
| ARC-AGI-2(high)[10] Compositional visual reasoning · ±9.2 is noise | |
| LiveBench[11] Composite score across 7 domains · ±2.7 is noise | |
| HMMTcontaminated[12] Competition mathematics — not ranked at any gap size | |
| AnalystAgent[13] Spreadsheet & document analysis · ±11.2 is noise |
Who ran these numbers: 12 of 13 independent — vals.ai (3), artificialanalysis.ai (2), tbench.ai (1), deepswe.datacurve.ai (1), toolathlon.xyz (1), snorkel.ai (1), arcprize.org (1), livebench.ai (1), matharena.ai (1); vendor self-reported (1).
- HLE: AA's own run, 'Claude Opus 4.8 (max)' — text-only, no tools; the board's underlying value is 48.66. Replaces the self-reported 49.8 from DeepSeek's comparison chart. Distinct from 'Fable 5 (Opus 4.8 fallback)' at 55.5 — a different system. The URL carries ?models= because AA's default board view no longer lists Opus 4.8 now that it is superseded: on the bare URL this row looks unsourced, which is exactly what scripts/check-sources.mjs flagged on 2026-08-19. Value re-confirmed on the filtered board the same day.
- HLE: As published in DeepSeek's comparison chart, not independently verified against Anthropic's own materials.
- Terminal-Bench 2.1: Re-verified 2026-09-03 in a browser after npm run check-sources flagged this row: tbench.ai shipped Terminal-Bench 4.0 as the new default board around 2026-08-27 and rebuilt the site so version is a client-side selection — the old deep link now 308s to a URL whose server-rendered HTML always carries the 4.0 board, and the 2.1 board only loads after picking "2.1" from the on-page selector, which is why the static check reads MISS. Value confirmed unchanged: 78.9%, rank 5 of 17, 'Opus 4.8 (high)' via Claude Code (CI now shown as ±2.6%, not ±1.3 — recomputed in the redesign, not a re-run). Vendor chart claimed 85.0. AA's own different-harness run gives 84.6 — harness choice alone spans 6 points here. Cross-check (2026-10-01): vals.ai's archived Terminus 2 table lists two 'Claude Opus 4.8' rows, 71.91 (priced $5/$25) and 69.66 (no price shown).
- DeepSWE: 59%±2 pass@1, mini-swe-agent harness, rank 10 (board updated 2026-08-13). Replaces the self-reported 58.0 — the two agree within error.
- Toolathlon-Verified: Label correction, same value: the board's 76.2±3.4 (rank 2) carries Toolathlon's 'Evaluated by us' badge — this was always an independent run, we had it mislabeled self-reported.
- Agents' Last Exam: Overall pass rate 27.0 (Claude Code, Max; score 45.1; $3,985). Replaces 25.7 — that value was the ALE-CLI split (and coincidentally the launch chart's claim); normalized to Overall + independent. Per-effort results on 2026-10-01: Max 27.0, XHigh 22.4, High 20.4.
- SWE-bench Verified: vals.ai run, rank 9/83 (updated 2026-08-14). The board lists a second cheaper-config Opus 4.8 row at 85.8; 88.6 is the primary entry.
- GPQA Diamond: vals.ai run, rank 14/133, displayed tied with DeepSeek V4 Pro 0813 (updated 2026-08-15).
- LiveCodeBench: vals.ai run, rank 9/138 (updated 2026-08-15).
- ARC-AGI-2: Claude Opus 4.8's official ARC-AGI-2 leaderboard row, dated 2026-06-01 on arcprize.org. The model's Max-tier cell is genuinely blank (N/A) for ARC-AGI-2, so High is its usable ceiling.
- LiveBench: Board row "Claude 4.8 Opus Thinking Max Effort" on the 2026-06-25 LiveBench release.
- HMMT: MathArena's leaderboard row is "Claude-Opus-4.8 (max)", 95.45% across all 33 HMMT Feb 2026 problems (live-verified 2026-08-25 at matharena.ai/?comp=hmmt--hmmt_feb_2026). Carries MathArena's own contamination-warning flag — "Model was released after competition release" — a disclosed risk that applies to every one of this site's 4 tracked HMMT Feb 2026 rows, Claude Opus 4.8 included, since all 4 tracked models were released after HMMT Feb 2026 took place.
- AnalystAgent: AA's own run, board row 'Claude Opus 4.8 (Adaptive Reasoning, Max Effort)'. pass@1 63.5, pass@5 78.75 — the widest ceiling-versus-reliability spread among the Claude rows.
Notes on the record
Superseded by Claude Opus 5 (2026-07-24, same $5/$25 pricing). Anthropic docs move it to 'Legacy models' with a migration guide, but it is formally still Active — tentative retirement not before 2027-05-28 — and serves as Claude Fable 5's automatic fallback for classifier-refused requests. Standard pricing shown; a $10/$50 'Fast Mode' tier is also offered. Confirmed current as of 2026-08-20: Anthropic's official Opus 5 announcement (https://www.anthropic.com/news/claude-opus-5) verifies the identical $5/$25 standard rate and the 2026-07-24 launch date.
Anthropic's refusals-and-fallback API docs (https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback) confirm Claude Opus 4.8's fallback role with a worked example: a response object shows a Claude Fable 5 refusal handed off with "from": claude-fable-5, "to": claude-opus-4-8. That same page states the category-to-fallback-model mapping used by default routing is not published, so which specific refusal categories route to Opus 4.8 cannot be confirmed from this source.
Fast Mode runs at up to 2.5x speed; Anthropic's 2026-05-28 launch post says this is three times cheaper than previous fast-mode pricing but does not name the prior price or generation — third-party pricing trackers cite roughly $30/$150 on the Opus 4.7 generation, but that figure is unconfirmed in Anthropic's own materials and should be verified before being treated as fact.
Compare with
FAQ
Has Claude Opus 4.8 been independently benchmarked?
Mostly yes. 12 of the 13 benchmark results tracked for Claude Opus 4.8 are independent third-party runs: Artificial Analysis (HLE without tools, AA-AnalystAgent), tbench.ai (Terminal-Bench 2.1), DataCurve (DeepSWE), toolathlon.xyz (Toolathlon-Verified), Snorkel (Agents' Last Exam), Vals AI (SWE-bench Verified, GPQA Diamond, LiveCodeBench), ARC Prize (ARC-AGI-2), LiveBench's own board, and MathArena (HMMT Feb 2026) (all observed 2026-08-17 through 2026-09-29). The one exception is HLE with tools (57.9, self-reported, observed 2026-08-13) — the 57.9 comes from DeepSeek's own launch comparison chart, not from Anthropic and not from an independent run — a competitor's claim about a rival's score, which is why it can never earn a Real-gap or Tie label here.
Is Claude Opus 4.8 still available now that Claude Opus 5 has launched?
Yes. Anthropic's docs have moved Claude Opus 4.8 into a 'Legacy models' section with a migration guide toward Opus 5, but its formal lifecycle status is still Active, with a tentative retirement no earlier than 2027-05-28. It also has a live operational role beyond the docs label: Anthropic's refusals-and-fallback API documentation uses Claude Opus 4.8 as the worked example for Claude Fable 5's classifier-refusal fallback routing, including a response example that hands a refused request off from claude-fable-5 to claude-opus-4-8.
How does Claude Opus 4.8 perform against Claude Opus 5 on the benchmarks this site tracks?
Opus 5 leads on four of the nine comparable shared benchmarks, ties on one, and four don't resolve to a clean call. On DeepSWE (74.0 vs 59.0, +15.0), Humanity's Last Exam without tools (54.9 vs 48.7, +6.2), LiveBench (80.1 vs 76.2, +3.9) and Agents' Last Exam (30.9 vs 27.0 at matched Max, +3.9), Opus 5 has four confirmed real gaps; on AA-AnalystAgent (53.75 vs 45.0 pass^5) the two are a statistical tie, inside the benchmark's noise band. Three more are graded saturated on this site — SWE-bench Verified, GPQA Diamond and, since 2026-09-29, LiveCodeBench (89.03 vs 87.82) — where both models score close to the ceiling and the test no longer separates frontier models, so none of those comparisons is meaningful regardless of the raw gap. The ninth, Terminal-Bench 2.1, is independently run on both sides but through different evaluation harnesses (tbench.ai's official board for Opus 4.8, Artificial Analysis for Opus 5), which this site treats as not a fair comparison. A tenth shared benchmark, ARC-AGI-2, doesn't yield a comparable line at all — the two models' tracked scores come from different reasoning-effort tiers (Opus 4.8's High-tier 72.1 vs Opus 5's Max-tier 90.4), so there's no matched-tier pair to call. At an identical $5/$25 list price, the comparisons that do resolve favor Opus 5 — the case for staying on Opus 4.8 is inertia (an existing integration already pointed at it), not price or measured capability. Separately, Opus 4.8 isn't going away regardless: Anthropic's own infrastructure still relies on it as Claude Fable 5's automatic fallback.
How does Claude Opus 4.8's pricing compare with Claude Opus 5's?
They're identical at standard rates: $5 per million input tokens and $25 per million output tokens for both. Opus 5 launched 2026-07-24 at parity with Opus 4.8 rather than at Fable 5's higher $10/$50 tier, per Anthropic's official Opus 5 announcement, which explicitly states its pricing is 'the same as Opus 4.8.' That means moving from Claude Opus 4.8 to Opus 5 costs nothing extra at list price — the case for switching rests on capability and Anthropic's own default routing, not budget.
What does Claude Opus 4.8's Fast Mode cost and what does it change?
Fast Mode runs Claude Opus 4.8 at up to 2.5x the base speed for $10 per million input tokens and $50 per million output tokens — a fixed premium over the $5/$25 standard rate. Anthropic's own launch announcement describes this as three times cheaper than previous fast-mode pricing, though it doesn't state the exact prior price or name which earlier generation it's compared against. Third-party pricing trackers cite roughly $30/$150 on the Opus 4.7 generation, but that specific figure isn't confirmed in Anthropic's own materials.
Why does Claude Opus 4.8 handle Claude Fable 5's refused requests?
When Claude Fable 5's built-in safety classifiers decline a request, Anthropic's API supports routing that refusal to another model automatically — either to whichever model Anthropic recommends for the refusal's category (server-side default routing) or to a fallback list you configure yourself. Anthropic's refusals-and-fallback API reference uses Claude Opus 4.8 as its worked example for both paths, including a response example that shows the handoff as 'from': claude-fable-5, 'to': claude-opus-4-8. The docs state the exact category-to-model mapping used by default routing isn't published, so this confirms Opus 4.8's documented fallback role without pinning down which specific refusal categories send traffic to it. That's a production dependency that exists independently of Opus 4.8's 'Legacy' label in the docs navigation.
Further reading
- GPQA Diamond leaderboard 2026 — Claude Opus 4.8 is one of the 23 models it compares.
- Models with 10M token context windows 2026 — Claude Opus 4.8 is one of the 34 models it compares.