Anthropic
Claude Fable 5.1
All six benchmark scores tracked here for Claude Fable 5.1 come from independent evaluators, not Anthropic's own runs — but that consensus is more fragile than it looks: Terminal-Bench 2.1 draws three different numbers for the identical model (91.4% from Artificial Analysis, 85.02% from vals.ai's own harness — its evaluation setup and tooling, no figure at all yet from the benchmark's canonical board tbench.ai), and vals.ai's own disclosed refusal-to-Opus fallback would cut its number further, to 79.03%, if counted as intended.
Claude Fable 5.1 benchmarks and pricing, every number sourced: 6 tracked Claude Fable 5.1 benchmark scores (6 independently run, 0 still resting on a vendor’s own claim), priced at $10.00 per million input tokens and $50.00 per million output.
Claude Fable 5.1’s 6 benchmark scores on this page were verified against their source on 2026-09-02.
Version history: succeeded Claude Fable 5 (2026-06-09).
- Released
- 2026-09-01
- License
- proprietary
- Context window
- 1M tokens
- Knowledge cutoff
- 2026-06
- Parameters
- Not disclosed
- Architecture
- Not disclosed
Claude Fable 5.1’s verified record
Against the 108 head-to-head comparisons Claude Fable 5.1 shares with other tracked models: 51 real gaps, 26 inside the noise band, and 31 we will not call.
A gap counts for Claude Fable 5.1 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Claude Fable 5.1 trails on 0 of them.
HLE · no tools Reasoning
±2 is noiseLiveBench Composite score across 7 domains
±2.7 is noiseLiveCodeBench Contest coding
±3.1 is noiseARC-AGI-2 · max Compositional visual reasoning
±9.2 is noiseTerminal-Bench 2.1 Terminal ops
±10.6 is noiseNo verdict for Claude Fable 5.1 anywhere on GPQA Diamond (saturated).
What changed from Claude Fable 5 to Claude Fable 5.1
The 6 benchmarks both models have been scored on, using the same variant each time. A raw Claude Fable 5.1 gain is not a real gain until it clears that benchmark’s own noise band, so each row below carries the verdict and not just the arithmetic.
| Benchmark | Claude Fable 5 | Claude Fable 5.1 | Change | Verdict |
|---|---|---|---|---|
| HLE | 55.5 | 59.1 | +3.6 | Real gap |
| Terminal-Bench 2.1 | 83.8 | 91.4 | +7.6 | Setup-dependent |
| GPQA Diamond | 93.18 | 93.43 | +0.3 | Tainted |
| LiveCodeBench | 89.78 | 90.52 | +0.7 | Tie |
| ARC-AGI-2 | 89.2 | 90 | +0.8 | Tie |
| LiveBench | 83 | 83.4 | +0.4 | Tie |
Claude Fable 5.1 API pricing
$10.00 in / $50.00 out per 1M tokens — official pricing
What Claude Fable 5.1 costs per job
| Workload | Tokens in / out | Cost |
|---|---|---|
| One long chat turn | 100K / 10K | $1.50 |
| A codebase review | 1,000K / 100K | $15.00 |
| A day of agent work | 10,000K / 1,000K | $150.00 |
Computed from Claude Fable 5.1’s list rates above — cache discounts and batch tiers are not applied.
Claude Fable 5.1 is one of 5 Anthropic models tracked on this site, at these official list prices.
| Model | In / 1M | Out / 1M |
|---|---|---|
| Claude Fable 5.1 | $10.00 | $50.00 |
| Claude Opus 5 | $5.00 | $25.00 |
| Claude Sonnet 5 | $2.00 | $10.00 |
| Claude Fable 5(superseded) | $10.00 | $50.00 |
| Claude Opus 4.8(superseded) | $5.00 | $25.00 |
Claude Fable 5.1 benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise | |
| Terminal-Bench 2.1[2] Terminal ops · ±10.6 is noise | |
| GPQA Diamondsaturated[3] Expert science Q&A — not ranked at any gap size | |
| LiveCodeBench[4] Contest coding · ±3.1 is noise | |
| ARC-AGI-2(max)[5] Compositional visual reasoning · ±9.2 is noise | |
| LiveBench[6] Composite score across 7 domains · ±2.7 is noise |
Who ran these numbers: 6 of 6 independent — artificialanalysis.ai (2), vals.ai (2), arcprize.org (1), livebench.ai (1).
- HLE: AA live leaderboard value, row labeled "Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)" — the board's top-ranked row for this model. AA's label doesn't name which model the fallback substitutes (Fable 5's own AA row is labeled "Opus 4.8 Fallback" by contrast); Anthropic names both Opus 5 and Opus 4.8 as fallback options without saying which applies when. Two lower AA effort tiers for the same model, Xhigh (58.7%) and High (55.9%), are on the same page but not recorded as separate rows, matching this site's one-row-per-model convention.
- Terminal-Bench 2.1: AA Max Effort row, rank 1 of 28 shown, Terminus 2 harness in an e2b sandbox. This benchmark shows three disagreeing numbers for this model — see "Notes on the record" on this page for the full comparison and why Artificial Analysis is the tracked figure.
- GPQA Diamond: vals.ai board, rank 8 of 137 systems, updated 2026-09-01. GPQA Diamond is graded saturated on this site, so this number never drives a verdict. Artificial Analysis separately reports 93.7% for the same model on its own harness — not recorded as a second row; both figures round to the same near-ceiling conclusion regardless of which is used.
- LiveCodeBench: vals.ai board, rank 1 of 142 systems, updated 2026-09-01.
- ARC-AGI-2: Top of five populated effort tiers (Max 90.0%, XHigh 90.0%, High 88.8%, Medium 86.3%, Low 78.3%); Max and XHigh tie on ARC-AGI-2, but ARC Prize's own listed inference cost is higher for Max ($4.49/task vs $3.12/task) — this is ARC Prize's compute cost to run the eval, not Anthropic's API price. Dated 2026-09-01 on arcprize.org. Companion score on the benchmark's earlier, easier predecessor, ARC-AGI-1 (not independently tracked on this site): 97.5%.
- LiveBench: Rank 1 of the entire board on the LiveBench-2026-06-25 release, Max Effort row. Category sub-scores: Reasoning 91.7, Coding 86.4, Agentic Coding 66.1, Mathematics 97.0, Data Analysis 80.3, Language 89.5, Instruction Following 73.0. LiveBench's own listed inference cost: $1.212 per successful task (its compute cost to run the eval, not Anthropic's API price).
Notes on the record
Claude Fable 5.1 launched September 1, 2026, per Anthropic's own announcement (independently corroborated by TechCrunch, Bloomberg, and VentureBeat coverage the same day; AWS's "What's New" post confirms the same day, "Posted on: Sep 1, 2026"). vals.ai's model page separately lists 2026-08-28, which reads as an internal API cutover ahead of the public announcement — this site uses the public GA date, matching how Claude Fable 5's own release_date uses its public launch rather than an earlier internal rollout. Anthropic's own models comparison table names it "Successor to Claude Fable 5" — the same relationship this site treats as a supersession elsewhere (Claude Opus 4.8 → Claude Opus 5, GLM-5.2 → GLM-5.3); Claude Fable 5's own entry on this site is marked superseded accordingly. Headline API pricing carries over unchanged at $10/MTok input and $50/MTok output, tied with Claude Fable 5 for the highest of the Claude models tracked here.
A less-restricted sibling exists, as with the prior generation: Claude Mythos 5.1 is the identical underlying model with different safety guardrails, limited to vetted organizations under Anthropic's invitation-only Project Glasswing. The context window carries over at 1M tokens in; max output is 128K tokens, though Fable 5's own max output was never independently sourced on this site, so that figure isn't confirmed as unchanged, only as Fable 5.1's own spec. The reliable knowledge cutoff moves five months later than Fable 5's, from January 2026 to June 2026 — and lands one month later than Claude Opus 5's own May 2026 cutoff on the same comparison table (source: platform.claude.com/docs/en/about-claude/models/overview, checked 2026-09-02). The one price that did move is cached reads, cut 75% from $1/MTok to $0.25/MTok; Anthropic's own announcement estimates this saves roughly 25% on typical workloads and up to 45% on highly agentic ones (source: anthropic.com/claude-fable-and-mythos-5-1).
Fable 5's most consequential provenance caveat carries over largely unchanged: Fable 5.1's safety classifiers can still reroute a refused prompt to another model as an automatic fallback, and vals.ai discloses running Fable 5.1 with exactly that fallback (to Claude Opus 5 or Claude Opus 4.8) in place for several of its evaluations. Where vals.ai has published the correction, the effect is measurable, not cosmetic: its own Terminal-Bench 2.1 score for this model falls from 85.02% to 79.03% once fallback-assisted answers are recounted as failures instead of passes. Anthropic's own release notes claim roughly 60% fewer cybersecurity refusal false positives from Fable 5.1's new safeguards versus Fable 5's old ones; a separate biology-safeguard update, which now applies to both Fable 5.1 and Fable 5 equally, cuts false positives on elementary biology and medical questions by roughly 85% versus what shipped at Fable 5's original launch (source: anthropic.com/claude-fable-and-mythos-5-1) — real improvements, if accurate, but neither claim nor the scores on this page settle how often the fallback still fires in practice.
Terminal-Bench 2.1 shows the starkest of these disagreements — three different pictures of the same model. Artificial Analysis's own harness — its evaluation setup and tooling, in this case Terminus 2 run in an e2b sandbox — reports 91.4%. vals.ai's own harness reports 85.02% (79.03% with the fallback correction above) for the identical model. And the benchmark's canonical board, tbench.ai, has not added Claude Fable 5.1 at all as of this writing — its leaderboard still lists only the unversioned "Claude Fable 5," unchanged since before this model existed. (GPQA Diamond shows two independent numbers too — 93.43% from vals.ai, 93.7% from Artificial Analysis — but that gap is small enough not to change any near-ceiling conclusion.) This site records the Artificial Analysis figure, following benchmarks.json's preferred-source order for this benchmark (tbench.ai first, Artificial Analysis second) now that tbench.ai has nothing to report; the vals.ai number is not discarded, only unused for the tracked score.
Same-day-of-launch coverage means several boards this site otherwise tracks have nothing yet: DeepSWE, Toolathlon-Verified, Agents' Last Exam, SWE-bench Verified, and HMMT Feb 2026 had not added Claude Fable 5.1 as of 2026-09-02, checked live on each. All six scores that do exist here — Humanity's Last Exam, Terminal-Bench 2.1, GPQA Diamond, LiveCodeBench, ARC-AGI-2, and LiveBench — come from independent evaluators (Artificial Analysis, vals.ai, ARC Prize, LiveBench's own board), none from Anthropic's own claimed numbers.
Compare with
FAQ
Has Claude Fable 5.1 been independently benchmarked?
Yes, on every benchmark this site has scored it against so far: all six tracked scores — Humanity's Last Exam, Terminal-Bench 2.1, GPQA Diamond, LiveCodeBench, ARC-AGI-2, and LiveBench — come from independent evaluators (Artificial Analysis, vals.ai, ARC Prize, and LiveBench's own board), not Anthropic's own claimed numbers. Five other boards this site tracks — DeepSWE, Toolathlon-Verified, Agents' Last Exam, SWE-bench Verified, and HMMT Feb 2026 — had not added the model as of 2026-09-02, one day after its public launch.
Does Claude Fable 5.1 still substitute a different model's answer when it refuses a prompt?
Yes — the mechanism from Claude Fable 5 carries over unchanged: when Fable 5.1's safety classifiers refuse a prompt, Claude Opus 5 or Claude Opus 4.8 answers instead, and the substituted answer can still count toward a published benchmark score. vals.ai discloses using this fallback for several of its Fable 5.1 evaluations; where it has published the correction, the effect is measurable — its Terminal-Bench 2.1 score for this model falls from 85.02% to 79.03% once fallback-assisted answers are recounted as failures. Anthropic's own release notes claim roughly 60% fewer cybersecurity refusal false positives from Fable 5.1's new safeguards versus Fable 5's old ones — a separate biology-safeguard update, now applied to both models equally, cuts false positives on elementary biology and medical questions by roughly 85% versus Fable 5's original launch. Both would reduce how often the fallback fires without eliminating it.
How does Claude Fable 5.1 pricing compare to Claude Fable 5?
The headline rate is identical: $10 per million input tokens and $50 per million output tokens, the same price Fable 5 has carried since its June 2026 launch, and the two of them remain tied for the highest of the Claude models tracked on this site. The one change is on cached reads, which drop 75% — from $1/MTok under Fable 5 to $0.25/MTok under Fable 5.1. Batch API pricing (50% off standard rates) is unchanged (source: platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1).
Is Claude Fable 5.1 included with a Claude subscription plan?
For chat use, yes: Anthropic lists Claude Fable 5.1 as available to Pro, Max, Team, and Enterprise subscribers (the Free plan is not included). Anthropic's pages did not state specific per-plan usage caps for Fable 5.1 at launch, pointing users to its general pricing page instead — Claude Fable 5's own record on this site shows those specifics can arrive and change weeks after a model's launch, so treat any cap figure elsewhere as provisional until Anthropic publishes one for 5.1 specifically (source: anthropic.com/claude/fable, checked 2026-09-02).
Does Claude Fable 5.1 replace Claude Fable 5, or sit alongside it?
It replaces it. Anthropic's own models comparison table describes Claude Fable 5.1 as the "successor to Claude Fable 5," the same relationship this site uses to mark a model superseded (as with Claude Opus 4.8 → Claude Opus 5 and GLM-5.2 → GLM-5.3) — Fable 5 stays on this site for comparison, not deleted, but is no longer the current recommendation. As with the earlier model, a less-restricted sibling exists: Claude Mythos 5.1 shares the same underlying weights and pricing but is limited to vetted organizations under Anthropic's invitation-only Project Glasswing. (For the full per-benchmark comparison against Fable 5, see the "What changed" table above.)
Further reading
- Models with 10M token context windows 2026 — Claude Fable 5.1 is one of the 23 models it compares.