OpenAI

GPT-6 Sol

Half GPT-5.6 Sol's promotional price at $2/$10 per million tokens, but no measurable gain over it on independent data: Humanity's Last Exam (47.9% vs 49.5%) and LiveBench (79.3 vs 81.0) both land inside their noise bands, same evaluator and max effort. AA's own Codex-agent DeepSWE run (69.0%) uses a different harness from GPT-5.6 Sol's board score. Eight of the eleven boards hadn't scored GPT-6 Sol yet.

GPT-6 Sol benchmarks and pricing, every number sourced: 3 tracked GPT-6 Sol benchmark scores (3 independently run, 0 still resting on a vendor’s own claim), priced at $2.00 per million input tokens and $10.00 per million output.

GPT-6 Sol’s 3 benchmark scores on this page were verified against their source on 2026-09-23.

Released
2026-09-22
License
proprietary
Context window
1M tokens
Knowledge cutoff
2026-04-20
Parameters
Not disclosed
Architecture
Not disclosed

GPT-6 Sol’s verified record

GPT-6 Sol’s most-compared rival is Claude Sonnet 5: 2 leads and 1 not callable across their 3 shared comparisons. GPT-6 Sol is priced at $2.00/$10.00 per 1M tokens (in/out) vs Claude Sonnet 5’s $2.00/$10.00.

Against the 77 head-to-head comparisons GPT-6 Sol shares with other tracked models: 31 real gaps, 20 inside the noise band, and 26 we will not call.

A gap counts for GPT-6 Sol only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — GPT-6 Sol trails on 8 of them.

HLE · no tools Reasoning

±2 is noise
Behind
GPT-6 Astra −6.8 · Claude Opus 5 −7.0 · Claude Fable 5.1 −11.2 · Claude Opus 5.5 −13.5 · 1 superseded: Claude Fable 5 −7.6
Ahead
Grok 4.7 +4.8 · Qwen3.8-Max +4.9 · Grok 4.6 +5.0 · GLM-5.3 +5.6 · Claude Sonnet 5 +6.6 · DeepSeek V4 Pro (0813) +6.9 · GLM-5.3-Flash +8.0 · GPT-5.6 Luna +8.4 · GPT-6 Luna +9.4 · Qwen3.8-Flash-Next +9.9 · 3 superseded: Muse Spark 1.2 +2.4 · GLM-5.2 +6.8 · DeepSeek V4 Flash (0731) +9.3
Tie
9 models within ±2
Unverified
1 model — vendor-reported on one side

LiveBench Composite score across 7 domains

±2.7 is noise
Behind
Claude Opus 5.5 −3.9 · Claude Fable 5.1 −4.1 · 1 superseded: Claude Fable 5 −3.7
Ahead
Qwen3.8-Flash-Next +3.1 · GLM-5.3 +3.2 · Claude Sonnet 5 +3.3 · GPT-5.6 Luna +5.7 · GPT-6 Luna +7.3 · GLM-5.3-Flash +7.7 · MiniMax M3 +12.0 · 3 superseded: Claude Opus 4.8 +3.1 · DeepSeek V4 Flash (0731) +5.1 · GLM-5.2 +6.1
Tie
11 models within ±2.7

No verdict for GPT-6 Sol anywhere on DeepSWE (nothing independently confirmed on both sides).

  • None of GPT-6 Sol’s coding comparisons are independently confirmed on both sides yet.

GPT-6 Sol API pricing

$2.00 in / $10.00 out per 1M tokens — official pricing

What GPT-6 Sol costs per job

GPT-6 Sol cost for three reference workloads, computed from its list rates
WorkloadTokens in / outCost
One long chat turn100K / 10K$0.300
A codebase review1,000K / 100K$3.00
A day of agent work10,000K / 1,000K$30.00

Computed from GPT-6 Sol’s list rates above — cache discounts and batch tiers are not applied.

GPT-6 Sol is one of 5 OpenAI models tracked on this site, at these official list prices.

OpenAI model pricing, official list rates
ModelIn / 1MOut / 1M
GPT-6 Luna$0.10$0.50
GPT-6 Sol$2.00$10.00
GPT-6 Astra$10.00$50.00
GPT-5.6 Luna$0.20$1.20
GPT-5.6 Sol$4.00$20.00

GPT-6 Sol benchmark scores

GPT-6 Sol benchmark scores, provenance, and source links
BenchmarkScore
HLE(no tools)[1]
Reasoning · ±2 is noise
DeepSWE[2]
Long-horizon coding · ±9.5 is noise
LiveBench[3]
Composite score across 7 domains · ±2.7 is noise

Who ran these numbers: 3 of 3 independent — artificialanalysis.ai (2), livebench.ai (1).

  1. HLE: AA's own run of GPT-6 Sol at max effort, text-only 2,158-question subset.
  2. DeepSWE: AA's own DeepSWE v1.1 run via its Codex agent harness (max effort) — part of AA's Coding Agent Index, not deepswe.datacurve.ai's own board, which does not list this model yet. OpenAI's self-reported figure is 68.8% (also max effort), within a point of AA's independent run.
  3. LiveBench: Board row 'GPT-6 Sol Max Effort' on the 2026-06-25 LiveBench release.

Notes on the record

Pricing is tiered by prompt length, confirmed on developers.openai.com/api/docs/models/gpt-6-sol: $2 input / $10 output per million tokens under 272,000 prompt tokens (cached input $0.20), rising to $4/$15 (cached $0.40) — 2x input and cache, 1.5x output, billed on the full request — once a prompt crosses that threshold. OpenAI frames this as a standard price, not a promotion: its own comparison table reads "$4→$2 input, $20→$10 output, 50% cheaper" against GPT-5.6 Sol, but explicitly calls GPT-5.6 Sol's $4/$20 rate "promotional pricing," so the headline "50% cheaper" compares against a temporary price, not GPT-5.6 Sol's original one.

Context window 1,050,000 tokens (max input 922,000), max output 128,000. Input: text and images. Output: text only. Knowledge cutoff April 20, 2026. License not stated for GPT-6 Sol; inferred proprietary, API-only.

No OpenAI page calls this a replacement for GPT-5.6 Sol — the announcement uses "predecessor" and "counterparts" when comparing scores, never "supersede" or "replace," and GPT-5.6 Sol carries no deprecation notice.

AA's own Codex-agent run of DeepSWE v1.1 (69.0%, max effort) is close to OpenAI's self-reported 68.8% (also max effort) — independent and vendor figures agree within a point. OpenAI also self-reports 56.4% on Agents' Last Exam, but names no split; this site's tracked variant is the Overall/full split, and without that confirmation the figure is not recorded here.

Compare with

FAQ

Has GPT-6 Sol been independently benchmarked?

Partly, as of one day post-launch — three of the eleven benchmarks this site tracks carry an independently-run score: Humanity's Last Exam (47.9%, Artificial Analysis, max effort), DeepSWE v1.1 via AA's own Codex-agent run (69.0%, max effort), and LiveBench (79.3 overall, max effort). GPQA Diamond, SWE-bench Verified, LiveCodeBench, Terminal-Bench 2.1, ARC-AGI-2, Toolathlon-Verified, Agents' Last Exam and HMMT Feb 2026 have none yet — deepswe.datacurve.ai, arcprize.org, tbench.ai, toolathlon.xyz, snorkel.ai, matharena.ai and vals.ai were all checked directly.

Is GPT-6 Sol better than GPT-5.6 Sol?

Not on anything this site can measure independently yet. On the two benchmarks where both have a score from the same evaluator at the same max effort, GPT-6 Sol comes out slightly lower — Humanity's Last Exam 47.9% versus 49.49%, LiveBench 79.3 versus 81.0 — and both gaps sit inside the benchmark's noise band, so they read as ties, not a regression. OpenAI's improvement claims rest on evaluations this site doesn't track, including its internal factuality test. What clearly changed is price: $2/$10 versus $4/$20.

Is GPT-6 Sol's price really 50% lower than GPT-5.6 Sol's?

Only against GPT-5.6 Sol's temporary rate. OpenAI's own comparison table shows $4→$2 input and $20→$10 output, but the same announcement calls GPT-5.6 Sol's $4/$20 rate "promotional pricing." GPT-6 Sol's $2/$10 is presented as OpenAI's new standard price, not a further promotion — so the "50% cheaper" framing compares against a price that was itself already discounted.

Does GPT-6 Sol's price change with longer prompts?

Yes. The $2 input / $10 output per-million-token rate applies to prompts up to 272,000 tokens. Once a prompt exceeds that threshold, OpenAI bills the whole request at 2x the input and cached-input rate and 1.5x the output rate — $4 input / $0.40 cached / $15 output — per developers.openai.com, confirmed 2026-09-23.

Does GPT-6 Sol replace GPT-5.6 Sol?

Not according to OpenAI's own materials. The announcement calls GPT-5.6 Sol a "predecessor" when comparing scores, never a replaced model, and GPT-5.6 Sol carries no deprecation notice on OpenAI's API deprecations page. Both models remain listed and priced.

Further reading

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.