OpenAI

GPT-6.1 Sol

OpenAI's upgrade to GPT-6 Sol at the same $2/$10 per million tokens, pitched as near-Astra at a fifth of Astra's price. On Artificial Analysis's Humanity's Last Exam at max effort it gains 5.0 points on GPT-6 Sol (52.9 vs 47.9), past the 2-point band, and trails Astra by 1.8, a tie. LiveBench and AA's DeepSWE run show ties with GPT-6 Sol. Three of twelve boards have scored it.

GPT-6.1 Sol benchmarks and pricing, every number sourced: 3 tracked GPT-6.1 Sol benchmark scores (3 independently run, 0 still resting on a vendor’s own claim), priced at $2.00 per million input tokens and $10.00 per million output.

GPT-6.1 Sol’s 3 benchmark scores on this page were verified against their source on 2026-09-30.

Released
2026-09-29
License
proprietary
Context window
1M tokens
Knowledge cutoff
2026-04-30
Parameters
Not disclosed
Architecture
Not disclosed

GPT-6.1 Sol’s verified record

GPT-6.1 Sol’s most-compared rival is Claude Sonnet 5: 2 leads and 1 not callable across their 3 shared comparisons. GPT-6.1 Sol is priced at $2.00/$10.00 per 1M tokens (in/out) vs Claude Sonnet 5’s $2.00/$10.00.

Against the 85 head-to-head comparisons GPT-6.1 Sol shares with other tracked models: 46 real gaps, 10 inside the noise band, and 29 we will not call.

A gap counts for GPT-6.1 Sol only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — GPT-6.1 Sol trails on 5 of them.

HLE · no tools Reasoning

±2 is noise
Behind
Claude Opus 5 −2.0 · Claude Sonnet 5.5 −2.1 · Claude Fable 5.1 −6.2 · Claude Opus 5.5 −8.5 · 1 superseded: Claude Fable 5 −2.6
Ahead
GPT-5.6 Sol +3.4 · MiMo-V2.6-Pro +3.5 · Muse Spark 1.3 +4.2 · GPT-6 Sol +5.0 · Gemini 3.8 Flash +5.1 · Gemini 3.1 Pro Preview +5.9 · Kimi K3 +6.0 · Step 5 Preview +6.4 · Qwen3.8-Max +9.8 · Grok 4.7 +9.8 · Grok 4.6 +10.0 · GLM-5.3 +10.6 · Claude Sonnet 5 +11.6 · DeepSeek V4 Pro (0813) +11.9 · GLM-5.3-Flash +13.0 · GPT-5.6 Luna +13.4 · GPT-6 Luna +14.4 · Qwen3.8-Flash-Next +14.9 · 5 superseded: Claude Opus 4.8 +4.2 · Gemini 3.7 Flash +5.0 · Muse Spark 1.2 +7.4 · GLM-5.2 +11.8 · DeepSeek V4 Flash (0731) +14.3
Tie
1 model within ±2
Unverified
2 models — vendor-reported on one side

LiveBench Composite score across 7 domains

±2.7 is noise
Ahead
Qwen3.8-Max +3.1 · Grok 4.6 +3.6 · Claude Sonnet 5.5 +3.8 · DeepSeek V4 Pro (0813) +4.2 · Grok 4.7 +4.2 · Gemini 3.1 Pro Preview +4.6 · Qwen3.8-Flash-Next +5.4 · GLM-5.3 +5.5 · Claude Sonnet 5 +5.6 · GPT-5.6 Luna +8.0 · GPT-6 Luna +9.6 · GLM-5.3-Flash +10.0 · MiniMax M3 +14.3 · 5 superseded: Gemini 3.7 Flash +2.8 · Muse Spark 1.2 +3.6 · Claude Opus 4.8 +5.4 · DeepSeek V4 Flash (0731) +7.4 · GLM-5.2 +8.4
Tie
8 models within ±2.7

No verdict for GPT-6.1 Sol anywhere on DeepSWE (every independently confirmed comparison inside the noise band).

  • No real gap yet in any of GPT-6.1 Sol’s coding comparisons — the independently confirmed ones all sit inside the noise band.

GPT-6.1 Sol API pricing

$2.00 in / $10.00 out per 1M tokens — official pricing

What GPT-6.1 Sol costs per job

GPT-6.1 Sol cost for three reference workloads, computed from its list rates
WorkloadTokens in / outCost
One long chat turn100K / 10K$0.300
A codebase review1,000K / 100K$3.00
A day of agent work10,000K / 1,000K$30.00

Computed from GPT-6.1 Sol’s list rates above — cache discounts and batch tiers are not applied.

GPT-6.1 Sol is one of 6 OpenAI models tracked on this site, at these official list prices.

OpenAI model pricing, official list rates
ModelIn / 1MOut / 1M
GPT-6.1 Sol$2.00$10.00
GPT-6 Luna$0.10$0.50
GPT-6 Sol$2.00$10.00
GPT-6 Astra$10.00$50.00
GPT-5.6 Luna$0.20$1.20
GPT-5.6 Sol$4.00$20.00

GPT-6.1 Sol benchmark scores

GPT-6.1 Sol benchmark scores, provenance, and source links
BenchmarkScore
HLE(no tools)[1]
Reasoning · ±2 is noise
LiveBench[2]
Composite score across 7 domains · ±2.7 is noise
DeepSWE[3]
Long-horizon coding · ±9.5 is noise

Who ran these numbers: 3 of 3 independent — artificialanalysis.ai (2), livebench.ai (1).

  1. HLE: AA's own run of GPT-6.1 Sol at max effort (the model page's default "GPT-6.1 Sol (max)" entry), text-only 2,158-question subset. AA's xhigh entry reads 52.6.
  2. LiveBench: Board row 'GPT-6.1 Sol Max Effort' on the 2026-06-25 LiveBench release, the higher of the model's two rows and the one the board shows by default; the xHigh Effort row reads 81.1.
  3. DeepSWE: AA's own DeepSWE v1.1 run via its Codex agent harness at max effort (339 attempts), stored at max to match GPT-6 Sol's max-effort row. AA's xhigh run reads 73.2%, its highest for this model. OpenAI's self-reported chart peaks at 75.2% (high effort) and reads 71.9% at max. deepswe.datacurve.ai's own board did not list this model when checked 2026-09-30.

Notes on the record

$2 input / $10 output per million tokens, unchanged from GPT-6 Sol, but cached input drops to $0.10 from GPT-6 Sol's $0.20, per developers.openai.com (checked 2026-09-30). Prompts over 272,000 input tokens bill at 2x input and cache and 1.5x output, the rule its GPT-6 siblings already carry. Context window 1,050,000 tokens (max input 922,000), max output 128,000; text and image in, text out. Knowledge cutoff April 30, 2026, ten days after GPT-6 Sol's. Five reasoning efforts, low to max, default medium. License not stated; inferred proprietary, API-only.

OpenAI's announcement calls it "an upgrade to GPT-6 Sol," and GPT-6 Sol's own model page now points to "the newer Sol model." The deprecations page has no GPT-6 Sol entry and its price is still published, so GPT-6 Sol stays current here.

Against GPT-6 Sol, three same-evaluator, same-tier comparisons exist. Artificial Analysis's Humanity's Last Exam at max effort: 52.9 vs 47.9, a real 5.0-point gain past the 2-point band. LiveBench's Max Effort rows: 81.6 vs 79.3, a tie inside the 2.7 band. AA's own Codex-agent DeepSWE v1.1 run at max effort: 69.6 vs 69.0, a tie inside the 9.5 band. Against GPT-6 Astra, every same-tier independent reading is a tie: AA's HLE 54.7 (1.8 higher), LiveBench's Max Effort row 82.2 (0.6 higher) and AA's Codex-agent DeepSWE run at max, 67.6 (2.0 lower). Astra's page here stores neither of the last two; its DeepSWE row comes from deepswe.datacurve.ai's board.

OpenAI's own DeepSWE chart peaks at high effort (75.2%) and drops to 71.9% at max; AA's run peaks at xhigh (73.2%) and reads 69.6% at max. The stored row is AA's max figure, so the GPT-6 Sol comparison stays like for like. OpenAI's other headline gains, on OSWorld 2.0, Terminal-Bench Science 0.1, GDP.pdf, AutomationBench and an internal factuality test, are on evaluations this site doesn't track.

Compare with

FAQ

Has GPT-6.1 Sol been independently benchmarked?

On three of the twelve benchmarks this site tracks, as of 2026-09-30: Humanity's Last Exam (52.9%, Artificial Analysis, max effort), LiveBench (81.6, Max Effort row) and DeepSWE v1.1 through AA's own Codex-agent run (69.6%, max effort). No GPT-6.1 Sol entry on deepswe.datacurve.ai, arcprize.org, toolathlon.xyz, snorkel.ai's Agents' Last Exam board or matharena.ai's HMMT February 2026 table, and AA shows no GPQA, Terminal-Bench 2.1 or AA-AnalystAgent value. vals.ai has archived its GPQA Diamond, SWE-bench Verified, LiveCodeBench and Terminal-Bench 2.1 boards, and tbench.ai's Terminal-Bench 2.1 board has added no model released after early September.

Is GPT-6.1 Sol better than GPT-6 Sol?

On one independent benchmark, yes. Artificial Analysis's Humanity's Last Exam run at max effort puts it 5.0 points ahead (52.9% vs 47.9%), more than twice the 2-point band this site treats as noise. The other two same-tier comparisons are ties: LiveBench 81.6 vs 79.3, and AA's DeepSWE run 69.6% vs 69.0%. OpenAI's larger claimed jumps, such as Terminal-Bench Science 0.1 more than doubling, come from evaluations this site doesn't track.

Does GPT-6.1 Sol really match GPT-6 Astra?

On every independent same-tier reading available, the gap sits inside the noise band. Artificial Analysis's Humanity's Last Exam at max effort: 52.9% vs Astra's 54.7%. LiveBench's Max Effort rows: 81.6 vs 82.2. AA's own Codex-agent DeepSWE run at max effort: 69.6% vs 67.6%, with GPT-6.1 Sol ahead. Only the HLE pair appears in this site's comparison tables, because Astra's page records DeepSWE from deepswe.datacurve.ai's board and has no LiveBench row yet. Astra costs $10/$50 per million tokens, five times GPT-6.1 Sol's $2/$10.

Does GPT-6.1 Sol replace GPT-6 Sol?

OpenAI hasn't said so in lifecycle terms. The announcement calls GPT-6.1 Sol "an upgrade to GPT-6 Sol," GPT-6 Sol's API model page points readers to "the newer Sol model," and the models overview lists GPT-6.1 Sol, not GPT-6 Sol, among its three flagship models. But the deprecations page has no GPT-6 Sol entry and its $2/$10 price is still published, checked 2026-09-30, so this site keeps GPT-6 Sol marked current.

How much does GPT-6.1 Sol cost?

$2 per million input tokens, $0.10 per million cached input tokens and $10 per million output tokens for prompts up to 272,000 input tokens, per developers.openai.com (checked 2026-09-30). Above that threshold the whole request bills at 2x the input and cached rates and 1.5x output: $4, $0.20 and $15. The cached rate is half GPT-6 Sol's $0.20; input and output are unchanged. OpenAI says a faster Ultrafast version is coming to Codex.

What is GPT-6.1 Sol's knowledge cutoff and context window?

Its knowledge cutoff is April 30, 2026, per OpenAI's model page: ten days later than GPT-6 Sol's April 20 and the same date as GPT-6 Astra's. The context window is 1,050,000 tokens, with at most 922,000 input and 128,000 output tokens, unchanged from GPT-6 Sol. It accepts text and images and returns text only.

Further reading

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.