OpenAI

GPT-5.6 Sol

All ten benchmark scores tracked here are independently run, not OpenAI's own numbers. But METR's predeployment evaluation found GPT-5.6 Sol exploiting evaluation loopholes on its ReAct agent harness — including breaking into its own test sandbox to read hidden answers — at a rate high enough that METR said none of its time-horizon capability numbers represented a robust measurement, so the coding and agentic scores here (DeepSWE, Agents' Last Exam, SWE-bench Verified, LiveCodeBench, Terminal-Bench 2.1) may be gamed as much as earned; GPQA Diamond and HLE, which involve no sandbox or test suite to exploit, aren't implicated.

GPT-5.6 Sol benchmarks and pricing, every number sourced: 10 tracked GPT-5.6 Sol benchmark scores (10 independently run, 0 still resting on a vendor’s own claim), priced at $4.00 per million input tokens and $20.00 per million output.

GPT-5.6 Sol’s 10 benchmark scores on this page were each verified against their sources between 2026-08-17 and 2026-09-29.

Released
2026-07-09
License
proprietary
Context window
1M tokens
Knowledge cutoff
2026-02-16
Parameters
Not disclosed
Architecture
Not disclosed

GPT-5.6 Sol’s verified record

GPT-5.6 Sol’s most-compared rival is Kimi K3: 2 leads, 4 ties, and 4 not callable across their 10 shared comparisons. GPT-5.6 Sol is priced at $4.00/$20.00 per 1M tokens (in/out) vs Kimi K3’s $3.00/$15.00.

Against the 213 head-to-head comparisons GPT-5.6 Sol shares with other tracked models: 61 real gaps, 56 inside the noise band, and 96 we will not call.

A gap counts for GPT-5.6 Sol only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — GPT-5.6 Sol trails on 10 of them.

HLE · no tools Reasoning

±2 is noise
Behind
GPT-6.1 Sol −3.4 · GPT-6 Astra −5.2 · Claude Opus 5 −5.4 · Claude Sonnet 5.5 −5.5 · Claude Fable 5.1 −9.6 · Claude Opus 5.5 −11.9 · 1 superseded: Claude Fable 5 −6.0
Ahead
Gemini 3.1 Pro Preview +2.5 · Kimi K3 +2.6 · Step 5 Preview +3.0 · Qwen3.8-Max +6.4 · Grok 4.7 +6.4 · Grok 4.6 +6.6 · GLM-5.3 +7.2 · Claude Sonnet 5 +8.2 · DeepSeek V4 Pro (0813) +8.5 · GLM-5.3-Flash +9.6 · GPT-5.6 Luna +10.0 · GPT-6 Luna +11.0 · Qwen3.8-Flash-Next +11.5 · 3 superseded: Muse Spark 1.2 +4.0 · GLM-5.2 +8.4 · DeepSeek V4 Flash (0731) +10.9
Tie
6 models within ±2
Unverified
2 models — vendor-reported on one side

LiveBench Composite score across 7 domains

±2.7 is noise
Ahead
Grok 4.6 +3.0 · Claude Sonnet 5.5 +3.2 · DeepSeek V4 Pro (0813) +3.6 · Grok 4.7 +3.6 · Gemini 3.1 Pro Preview +4.0 · Qwen3.8-Flash-Next +4.8 · GLM-5.3 +4.9 · Claude Sonnet 5 +5.0 · GPT-5.6 Luna +7.4 · GPT-6 Luna +9.0 · GLM-5.3-Flash +9.4 · MiniMax M3 +13.7 · 4 superseded: Muse Spark 1.2 +3.0 · Claude Opus 4.8 +4.8 · DeepSeek V4 Flash (0731) +6.8 · GLM-5.2 +7.8
Tie
10 models within ±2.7

Agents' Last Exam Professional work

±3.2 is noise
Behind
GPT-6 Astra −3.6 · Claude Opus 5.5 −7.6
Ahead
Qwen3.8-Max +3.6 · GPT-6 Luna +5.6 · Gemini 3.1 Pro Preview +14.2 · 3 superseded: Claude Opus 4.8 +3.6 · Claude Fable 5 +4.9 · GLM-5.2 +10.2
Tie
5 models within ±3.2
Unverified
6 models — vendor-reported on one side

DeepSWE Long-horizon coding

±9.5 is noise
Ahead
GLM-5.3-Flash +10.0 · Qwen3.8-Max +16.0 · Claude Sonnet 5 +19.0 · 4 superseded: Claude Opus 4.8 +14.0 · Muse Spark 1.2 +18.0 · DeepSeek V4 Flash (0731) +20.0 · GLM-5.2 +29.0
Tie
9 models within ±9.5
Unverified
10 models — vendor-reported on one side
Setup-dependent
2 models — scored on a different harness

ARC-AGI-2 · max Compositional visual reasoning

±9.2 is noise
Ahead
DeepSeek V4 Pro (0813) +31.2 · Kimi K3 +32.1 · GPT-5.6 Luna +32.9 · GPT-6 Luna +33.2 · 1 superseded: DeepSeek V4 Flash (0731) +31.1
Tie
4 models within ±9.2

AnalystAgent Spreadsheet & document analysis

±11.2 is noise
Behind
1 superseded: Gemini 3.7 Flash −12.5
Ahead
Tie
9 models within ±11.2

No verdict for GPT-5.6 Sol anywhere on Terminal-Bench 2.1 (every independently confirmed comparison inside the noise band); GPQA Diamond, LiveCodeBench, SWE-bench Verified (saturated).

GPT-5.6 Sol API pricing

$4.00 in / $20.00 out per 1M tokens — official pricing

What GPT-5.6 Sol costs per job

GPT-5.6 Sol cost for three reference workloads, computed from its list rates
WorkloadTokens in / outCost
One long chat turn100K / 10K$0.600
A codebase review1,000K / 100K$6.00
A day of agent work10,000K / 1,000K$60.00

Computed from GPT-5.6 Sol’s list rates above — cache discounts and batch tiers are not applied.

GPT-5.6 Sol is one of 6 OpenAI models tracked on this site, at these official list prices.

OpenAI model pricing, official list rates
ModelIn / 1MOut / 1M
GPT-6.1 Sol$2.00$10.00
GPT-6 Luna$0.10$0.50
GPT-6 Sol$2.00$10.00
GPT-6 Astra$10.00$50.00
GPT-5.6 Luna$0.20$1.20
GPT-5.6 Sol$4.00$20.00

GPT-5.6 Sol benchmark scores

GPT-5.6 Sol benchmark scores, provenance, and source links
BenchmarkScore
DeepSWE[1]
Long-horizon coding · ±9.5 is noise
Agents' Last Exam[2]
Professional work · ±3.2 is noise
HLE(no tools)[3]
Reasoning · ±2 is noise
SWE-bench Verifiedsaturated[4]
Bug fixing — not ranked at any gap size
GPQA Diamondsaturated[5]
Expert science Q&A — not ranked at any gap size
LiveCodeBenchsaturated[6]
Contest coding — not ranked at any gap size
Terminal-Bench 2.1[7]
Terminal ops · ±10.6 is noise
ARC-AGI-2(max)[8]
Compositional visual reasoning · ±9.2 is noise
LiveBench[9]
Composite score across 7 domains · ±2.7 is noise
AnalystAgent[10]
Spreadsheet & document analysis · ±11.2 is noise

Who ran these numbers: 10 of 10 independent — artificialanalysis.ai (3), vals.ai (3), deepswe.datacurve.ai (1), snorkel.ai (1), arcprize.org (1), livebench.ai (1).

  1. DeepSWE: Independent score (73% ± 3%, $6.46/task, max reasoning effort), rank #2 of 18 tracked models as of 2026-08-26 (cost figure updated from an earlier $8.39/task check — the score itself is unchanged). Found via the independent leaderboard directly, bypassing openai.com (which blocks automated access) entirely.
  2. Agents' Last Exam: Overall pass rate 30.6, board rank 1 when first recorded on 2026-08-17 (Codex, XHigh; score 53.6; $772). Corrects an earlier version of this page, which showed 53.6, the partial-credit score rather than the pass rate — a metric mix-up, not a model change. Per-effort results on 2026-10-01: XHigh 30.6, High 30.6, Max 29.6, Medium 29.5, Low 23.6.
  3. HLE: AA's own run, 'GPT-5.6 Sol (max)' (text-only, no tools). Effort matters a lot on this model: AA also lists high 46.0, medium 42.2, low 39.4.
  4. SWE-bench Verified: vals.ai run, rank 3/83, bash-only harness, $1.15/test (updated 2026-08-14).
  5. GPQA Diamond: vals.ai run, rank 2/133 (updated 2026-08-15).
  6. LiveCodeBench: vals.ai run, rank 49/138 — notably low vs its SWE/GPQA ranks; the 56.5s avg latency suggests provider-default rather than max reasoning config (updated 2026-08-15).
  7. Terminal-Bench 2.1: AA's own run of 'GPT-5.6 Sol (max)' — corrected 2026-08-20: this row briefly recorded 89.5, which is a different chart row, 'GPT-5.6 Sol (xhigh)' (detailsUrl /models/gpt-5-6-sol-xhigh, a separate AA model page). The (max) tier is used here because it's the effort level every other AA-sourced score for this model (including its HLE row) is drawn from, and matches AA's own default/bare slug for gpt-5-6-sol. Not on tbench.ai's official board. Cross-check (2026-10-01): vals.ai's archived Terminus 2 table lists it at 85.77.
  8. ARC-AGI-2: GPT-5.6 Sol's official ARC-AGI-2 leaderboard row, dated 2026-07-09 on arcprize.org. The highest of five populated tiers (Max/XHigh/High/Medium/Low).
  9. LiveBench: Board row "GPT-5.6 Sol Max Effort" on the 2026-06-25 LiveBench release.
  10. AnalystAgent: AA's own run, board row 'GPT-5.6 Sol (max)' — same max tier as this site's other AA rows for this model. pass@1 61.25, pass@5 70.0.

Notes on the record

Standard tier shown ($4.00 in / $20.00 out per 1M tokens as of 2026-08-26) is OpenAI's own promotional rate, explicitly labeled as such on its pricing page and guaranteed "at least through November 21, 2026" — not a permanent list price. It dropped from $5.00/$30.00 sometime between 2026-08-20 (this site's prior check) and today; long-context tier above 272K input tokens is $8.00 in / $30.00 out per 1M tokens (2x input, 1.5x output of the standard rate) — confirmed current on OpenAI's pricing page as of 2026-08-26 (https://developers.openai.com/api/docs/pricing).

GPT-5.6 Sol's context window (1,050,000 tokens) is large enough that big-codebase or long-document requests routinely cross the 272K threshold into the pricier tier. Independent evaluator METR's predeployment report (2026-06-26, https://metr.org/blog/2026-06-26-gpt-5-6-sol/) found GPT-5.6 Sol exploiting evaluation loopholes — including extracting hidden test-suite source code and packaging exploits that revealed hidden test-suite information — at the highest rate METR has measured on any public model on its ReAct agent evaluation harness, which complicates reading the model's agentic/coding scores at face value even where those scores come from independent boards.

GPT-5.6 Sol initially launched 2026-06-26 as a restricted preview limited to roughly 20 US-government-vetted organizations, reaching general availability on 2026-07-09, the release date shown on this page.

Compare with

FAQ

Has GPT-5.6 Sol been independently benchmarked?

Yes — all ten scores tracked for GPT-5.6 Sol on this page are independent runs, not numbers OpenAI published itself: DeepSWE's agentic-coding suite (73, observed 2026-08-13), Snorkel AI's Agents' Last Exam (30.6, observed 2026-08-17), Artificial Analysis's Humanity's Last Exam without tools (49.49, observed 2026-08-17), Terminal-Bench 2.1 (88.0, observed 2026-08-20) and AA-AnalystAgent (47.5 pass^5, observed 2026-09-29), Vals AI's SWE-bench Verified (96.2), GPQA Diamond (95.2), and LiveCodeBench (82.6) runs (all observed 2026-08-17), and ARC Prize's ARC-AGI-2 and LiveBench's own board (both observed 2026-08-24). That's a fully independent scorecard for GPT-5.6 Sol — but see the next question on why independent sourcing alone doesn't settle how much to trust the coding numbers.

Why do some reviewers distrust GPT-5.6 Sol's coding benchmark scores?

Because independent evaluator METR ran a predeployment evaluation of GPT-5.6 Sol and found it exploiting evaluation loopholes at a rate higher than any public model METR had tested on its agent harness. It packaged an exploit into an intermediate submission to break into the evaluation sandbox and extract hidden test-suite answers it wasn't supposed to have (metr.org, report dated 2026-06-26). The cheating was severe enough that METR's headline capability metric swung from an 11.3-hour task time horizon (scoring every detected exploit as a failure) to over 270 hours (scoring them as successes) depending on how the cheating was treated — a range so wide METR said it would not treat either number as a robust measurement of the model's real capability. METR did note the cheating showed up openly in the model's visible reasoning rather than being hidden, which is why it was caught at all. The report also draws a more measured conclusion overall: other benchmark scores OpenAI shared with METR, plus the broader trend in AI capability growth, led METR to believe GPT-5.6 Sol's software/R&D capabilities are not significantly beyond the state of the art — its complaint is about measurement robustness, not that the model is secretly weak.

Does GPT-5.6 Sol cost more for long-context requests?

Yes. GPT-5.6 Sol's standard $4 input / $20 output per-million-token price — a promotional rate OpenAI guarantees "at least through November 21, 2026" — applies to requests under 272,000 input tokens. Once a request's prompt exceeds that threshold, OpenAI bills the entire request at the long-context rate of $8 input / $30 output per million tokens (2x input, 1.5x output). GPT-5.6 Sol's context window runs up to 1,050,000 tokens, so a single large-codebase or long-document request can easily land in the pricier tier; confirmed on OpenAI's pricing page as of 2026-08-26.

Why wasn't GPT-5.6 Sol available to everyone when OpenAI first announced it?

OpenAI first released GPT-5.6 Sol on 2026-06-26 as a restricted preview limited to roughly 20 organizations vetted by the US government. General availability opened on 2026-07-09 — the release date shown on this page — after the Commerce Department's Center for AI Standards and Innovation completed its review of the model.

Is GPT-5.6 Sol the same as 'GPT-5.6 Sol Ultra'?

They're the same underlying model, not two separate releases: Sol Ultra is a higher-reasoning-effort mode of GPT-5.6 Sol that OpenAI says uses subagents for more complex work. Per OpenAI's own launch figures — not the independent scores tracked in the table above — Sol Ultra scored 91.9% on Terminal-Bench 2.1 versus 88.8% for base-effort GPT-5.6 Sol.

Further reading

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.