Anthropic
Claude Haiku 5.5
Anthropic's cheapest and fastest small model: $0.10/$0.50 per million tokens up to a 100,000-token prompt, then $0.50/$2.50 — the lowest long-context threshold on this site. Three boards scored it within a day: HLE 44.4 (Artificial Analysis, max effort), LiveBench 72.1 (default xHigh row; Max reads 67.9) and Terminal-Bench 4.0 35.4 (vals.ai, max effort), enough to sit above GPT-6 Luna there. Only three of its five scores are independent.
Claude Haiku 5.5 benchmarks and pricing, every number sourced: 5 tracked Claude Haiku 5.5 benchmark scores (3 independently run, 2 still resting on a vendor’s own claim), priced at $0.10 per million input tokens and $0.50 per million output.
- Released
- 2026-10-07
- License
- proprietary
- Context window
- 1M tokens
- Knowledge cutoff
- 2026-06
- Verified
- source checked 2026-10-08
- Parameters
- Not disclosed
- Architecture
- Not disclosed
Claude Haiku 5.5’s verified record
Claude Haiku 5.5’s most-compared rival is Claude Sonnet 5.5: 3 trails and 1 not callable across their 4 shared comparisons. Claude Haiku 5.5 is priced at $0.10/$0.50 per 1M tokens (in/out) vs Claude Sonnet 5.5’s $2.00/$10.00.
Against the 91 head-to-head comparisons Claude Haiku 5.5 shares with other tracked models: 49 real gaps, 10 inside the noise band, and 32 we will not call.
A gap counts for Claude Haiku 5.5 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Claude Haiku 5.5 trails on 32 of them.
HLE · no tools Reasoning
±2 is noiseLiveBench Composite score across 7 domains
±2.7 is noiseTerminal-Bench 4.0 · max Terminal ops
±12.4 is noiseNo verdict for Claude Haiku 5.5 anywhere on HLE · with tools ( nothing independently confirmed on both sides); OSWorld 2.0 · v2 1 offline ( every comparison inside the noise band).
Claude Haiku 5.5 API pricing
$0.10 in / $0.50 out per 1M tokens — official pricing source
What Claude Haiku 5.5 costs per job
| Workload | Tokens in / out | Cost |
|---|---|---|
| One long chat turn | 100K / 10K | $0.015 |
| A codebase review | 1,000K / 100K | $0.150 |
| A day of agent work | 10,000K / 1,000K | $1.50 |
Computed from Claude Haiku 5.5’s list rates above — cache discounts and batch tiers are not applied.
Claude Haiku 5.5 is one of 8 Anthropic models tracked on this site, at these official list prices.
| Model | In / 1M | Out / 1M |
|---|---|---|
| Claude Haiku 5.5 | $0.10 | $0.50 |
| Claude Sonnet 5.5 | $2.00 | $10.00 |
| Claude Opus 5.5 | $4.00 | $20.00 |
| Claude Fable 5.1 | $10.00 | $50.00 |
| Claude Opus 5(previous version) | $5.00 | $25.00 |
| Claude Sonnet 5(previous version) | $2.00 | $10.00 |
| Claude Fable 5(previous version) | $10.00 | $50.00 |
| Claude Opus 4.8(previous version) | $5.00 | $25.00 |
Claude Haiku 5.5 benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise | |
| HLE(with tools)[2] Reasoning · ±2 is noise | |
| LiveBench[3] Composite score across 7 domains · ±2.7 is noise | |
| Terminal-Bench 4.0(max)[4] Terminal ops · ±12.4 is noise | |
| OSWorld 2.0(v2 1_offline)[5] Long-horizon computer use · ±9.7 is noise |
Who ran these numbers: 3 of 5 independent — artificialanalysis.ai (1), livebench.ai (1), vals.ai (1); vendor self-reported (2).
- HLE: Artificial Analysis's own run, no tools: their model page titles the row Max and the embedded data tags effort max, so max is their tier, not an inference. Anthropic's own no-tools figure is 45.9 — 1.5 higher — run with Claude Opus 4.6 as grader, thinking on auto and a 980k-token task budget, no context compaction (system card §8.8.1). Read the two as one independent and one vendor run of the same benchmark, not as a disagreement about the model.
- HLE: Anthropic launch chart (2026-10-07), no effort tier printed on the row — the effort field here is undeclared, not max. System card §8.8.1 gives the configuration: web search, web fetch, programmatic tool calling and code execution, thinking auto, 980k task budget, Claude Opus 4.6 as grader. Tools move this model 11.5 points on Anthropic's own table (57.4 vs 45.9 no tools), the same scaffold effect this site tracks elsewhere; GPT-6 Luna is blank in Anthropic's table.
- LiveBench: LiveBench's 2026-06-25 release, recomputed with this site's method (mean of the seven category averages): 72.07 from the board's 23 task columns. Two rows exist and the higher tier reads lower — xHigh 72.07, Max Effort 67.88 — so this site records the default-displayed xHigh row per its LiveBench convention and reports the Max figure here rather than silently picking the better number.
- Terminal-Bench 4.0: vals.ai's own run (their model card, checked 2026-10-08): 35.35 ± 2.20, rank 12 of 45. The row carries its own refusal tooltip — 3 provider refusals (1.51%), scored as failures — which replaces the card header's aggregate 0.00% fallback and 0.22% refusal figures; those describe the card's Vals Index composite, not this run. The card's single default hyperparameter block reads temperature 1, 128K max output and compute effort max, with vals' own warning that individual benchmarks may differ. Anthropic's own figure for the same benchmark is 39.2 with safeguards enabled and no fallback model, where 1.8% of trials (12 of 660, ten on one task) stopped on a safeguard and failed (system card §8.4) — a protocol difference, not a scoring dispute, so the vendor number stays in this note rather than becoming a row.
- OSWorld 2.0: Anthropic launch chart (2026-10-07), labelled OSWorld 2.1, partial reward on the offline subset; system card §8.9.3 confirms max effort with benchmark defaults (1080p, 500 steps, Opus 4.8 grader) and gives a strict pass rate of 37.1% alongside the 72.4 partial. Not comparable with the v2_1_full rows other Claude models carry on this benchmark (Sonnet 5.5 reads 80.1 on the full environment), because the offline subset is a different task set. No official-board row for this model: the board's published results file (updated 2026-10-05, checked 2026-10-08) lists 51 rows and none is Haiku.
Notes on the record
$0.10 in / $0.50 out per million tokens up to a 100,000-token prompt, then $0.50 / $2.50 (Anthropic's pricing table, checked 2026-10-08) — a 5x step at a lower threshold than anywhere else here, since Sonnet 5 and Opus 5 switch at 272K. Anthropic puts the average saving against Haiku 4.5 at about 75%. Knowledge cutoff June 2026. Adaptive thinking is on by default at medium effort, the first Haiku with the effort parameter. Anthropic publishes no parameter count and no licence terms here, so the licence field is inferred from the closed API, not vendor-stated.
Three boards have scored it. Artificial Analysis's Humanity's Last Exam, no tools at max effort, reads 44.39; Anthropic's own no-tools figure is 45.9, 1.5 higher, graded by Claude Opus 4.6 on a 980k task budget. LiveBench's default xHigh row reads 72.07 overall against 67.88 at Max Effort — another tier inversion of the usual higher-tier-reads-higher order, the same shape as Opus 5.5's ARC-AGI-2 row. vals.ai's Terminal-Bench 4.0 run reads 35.35 ± 2.20 at max effort, its row reporting 3 provider refusals (1.51%) scored as failures, against Anthropic's own 39.2.
That last pair is not like-for-like. Anthropic ran Terminal-Bench 4.0 with no fallback model: safeguards stopped 1.8% of trials (12 of 660, ten of them on a single task) and every one of those trials failed. Sonnet 5.5's cyber and frontier-LLM safety blocks fall back to Sonnet 5 by default, so its rows can contain another model's answers; Haiku 5.5's cannot.
No row exists on GPQA Diamond, SWE-bench Verified, LiveCodeBench or Terminal-Bench 2.1: vals.ai archived those boards, and Artificial Analysis published none for models released after 2026-09-11. Agents' Last Exam, AA-AnalystAgent, HMMT Feb 2026, DeepSWE and Toolathlon-Verified show no row, each checked directly on 2026-10-08; ARC-AGI-2's board now redirects to Kaggle, so it is unchecked rather than confirmed clear. Three of the fourteen benchmarks here carry an independent number, a fourth (OSWorld 2.1) is vendor-only.
Compare with
FAQ
What does Claude Haiku 5.5 cost?
$0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, then $0.50 and $2.50 above that. Cache reads are $0.01 / $0.05, a 5-minute cache write $0.125 / $0.625, and the batch API is half price. Against Haiku 4.5's $1.00 / $5.00 that is 90% cheaper under the threshold and 50% cheaper over it. Weighting those by Anthropic's own 90/10 request split gives about 86%, and its headline 75% average saving additionally absorbs the newer tokenizer's roughly 30% extra tokens per task.
What has Claude Haiku 5.5 been benchmarked on?
Three independent scores, all found within a day of launch: Humanity's Last Exam 44.39 no tools (Artificial Analysis, max effort), LiveBench 72.07 overall (the board's default xHigh row; Max Effort reads 67.88, a tier inversion of the usual order) and Terminal-Bench 4.0 35.35 ± 2.20 (vals.ai, max effort). OSWorld 2.1 72.4 is Anthropic's own figure, on the offline subset. The other ten show no row as of 2026-10-08 — four because vals.ai archived those boards; ARC-AGI-2's board redirects to Kaggle and is unchecked.
Is Claude Haiku 5.5 better than Claude Haiku 4.5?
On Anthropic's own runs, by a wide margin: Humanity's Last Exam no tools 45.9 against 10.2, Terminal-Bench 4.0 39.2 against 0.0, and OSWorld 2.1 offline 72.4 against 15.7. All three are vendor-run, and this site has no independent Haiku 4.5 row to check them against. The sticker price is not the whole saving: it uses the newer tokenizer from Claude 4.7 on, so the same task costs about 30% more tokens than on 4.5, which narrows the gap between the two price points.
Is Claude Haiku 5.5 good enough for coding agents?
Not as the main model. On vals.ai's Terminal-Bench 4.0 board it reads 35.35 against 64.14 for Sonnet 5.5 and 65.15 for Opus 5.5 — real gaps on a 12.4-point noise band. On LiveBench it trails Sonnet 5.5's 77.8 by 5.7, a real gap past the 2.7 band at matched xHigh effort. Anthropic's own framing is narrower than agentic coding: compaction, summarisation, classification and subagent work, where 5% of Sonnet 5.5's price changes the arithmetic.
Why does Claude Haiku 5.5 get more expensive on long prompts?
Because the rate changes above 100,000 input tokens, from $0.10/$0.50 to $0.50/$2.50 — five times both. Anthropic says prompts under 100K are about 90% of requests to its previous Haiku, so most traffic never meets the step-up; past it, Haiku 5.5 costs more than GPT-6 Luna ($0.10/$0.50) and considerably more per output token than MiMo-V2.6-Flash ($0.14/$0.28).
What is the effort parameter on Claude Haiku 5.5?
Adaptive thinking is on by default, at medium effort, and this is the first Haiku-class model where the effort parameter is available at all. Thinking can be disabled at high effort or below. It matters when reading scores: on LiveBench the xHigh row (72.07) reads higher than the Max Effort row (67.88), so quoting one number without its tier hides a 4.2-point difference.
Further reading
- Models with 10M token context windows 2026 — Claude Haiku 5.5 is one of the 37 models it compares.
Benchmark guides
- Humanity's Last Exam — what Claude Haiku 5.5’s reasoning score on it does and doesn’t prove.
- LiveBench — what Claude Haiku 5.5’s composite score on it does and doesn’t prove.