OpenAI

GPT-6 Luna

The cheapest GPT-6 at $0.10/$0.50 per million tokens — half GPT-5.6 Luna's promotional input rate and 42% of its output rate — with no measurable gain over it on independent data: Humanity's Last Exam (38.5% vs 39.5%), LiveBench (72.0 vs 73.6) and ARC-AGI-2 (59.31% vs 59.6%) all tie inside their noise bands, same evaluators and max effort. Eight of the eleven boards hadn't scored GPT-6 Luna yet.

GPT-6 Luna benchmarks and pricing, every number sourced: 4 tracked GPT-6 Luna benchmark scores (3 independently run, 1 still resting on a vendor’s own claim), priced at $0.10 per million input tokens and $0.50 per million output.

GPT-6 Luna’s 4 benchmark scores on this page were each verified against their sources between 2026-09-22 and 2026-09-23.

Released
2026-09-22
License
proprietary
Context window
1M tokens
Knowledge cutoff
2026-05-18
Parameters
Not disclosed
Architecture
Not disclosed

GPT-6 Luna’s verified record

GPT-6 Luna’s most-compared rival is Claude Opus 5: 3 trails and 1 not callable across their 4 shared comparisons. GPT-6 Luna is priced at $0.10/$0.50 per 1M tokens (in/out) vs Claude Opus 5’s $5.00/$25.00.

Against the 86 head-to-head comparisons GPT-6 Luna shares with other tracked models: 48 real gaps, 12 inside the noise band, and 26 we will not call.

A gap counts for GPT-6 Luna only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — GPT-6 Luna trails on 47 of them.

HLE · no tools Reasoning

±2 is noise
Behind
DeepSeek V4 Pro (0813) 2.5 · Claude Sonnet 5 2.8 · GLM-5.3 3.8 · Grok 4.6 4.4 · Qwen3.8-Max 4.5 · Grok 4.7 4.6 · Step 5 Preview 8.0 · Kimi K3 8.4 · Gemini 3.1 Pro Preview 8.5 · Gemini 3.8 Flash 9.3 · GPT-6 Sol 9.4 · Muse Spark 1.3 10.6 · MiMo-V2.6-Pro 10.9 · GPT-5.6 Sol 11.0 · GPT-6 Astra 16.2 · Claude Opus 5 16.4 · Claude Fable 5.1 20.6 · Claude Opus 5.5 22.9 · 5 superseded: GLM-5.2 2.6 · Muse Spark 1.2 7.0 · Gemini 3.7 Flash 9.4 · Claude Opus 4.8 10.2 · Claude Fable 5 17.0
Tie
4 models within ±2
Unverified
1 model — vendor-reported on one side

LiveBench Composite score across 7 domains

±2.7 is noise
Behind
Claude Sonnet 5 4.0 · GLM-5.3 4.1 · Qwen3.8-Flash-Next 4.2 · Gemini 3.1 Pro Preview 5.0 · DeepSeek V4 Pro (0813) 5.4 · Grok 4.7 5.4 · Grok 4.6 6.0 · Qwen3.8-Max 6.5 · Kimi K3 7.2 · GPT-6 Sol 7.3 · Claude Opus 5 8.1 · GPT-5.6 Sol 9.0 · Muse Spark 1.3 9.6 · Claude Opus 5.5 11.2 · Claude Fable 5.1 11.4 · 4 superseded: Claude Opus 4.8 4.2 · Muse Spark 1.2 6.0 · Gemini 3.7 Flash 6.8 · Claude Fable 5 11.0
Ahead
Tie
4 models within ±2.7

ARC-AGI-2 · max Compositional visual reasoning

±9.2 is noise
Behind
Claude Fable 5.1 30.7 · Claude Opus 5 31.1 · GPT-5.6 Sol 33.2 · GPT-6 Astra 35.7 · 1 superseded: Claude Fable 5 29.9
Tie
4 models within ±9.2

No verdict for GPT-6 Luna anywhere on DeepSWE (nothing independently confirmed on both sides).

  • None of GPT-6 Luna’s coding comparisons are independently confirmed on both sides yet.

GPT-6 Luna API pricing

$0.10 in / $0.50 out per 1M tokens official pricing

What GPT-6 Luna costs per job

GPT-6 Luna cost for three reference workloads, computed from its list rates
WorkloadTokens in / outCost
One long chat turn100K / 10K$0.015
A codebase review1,000K / 100K$0.150
A day of agent work10,000K / 1,000K$1.50

Computed from GPT-6 Luna’s list rates above — cache discounts and batch tiers are not applied.

GPT-6 Luna is one of 5 OpenAI models tracked on this site, at these official list prices.

OpenAI model pricing, official list rates
ModelIn / 1MOut / 1M
GPT-6 Luna$0.10$0.50
GPT-6 Sol$2.00$10.00
GPT-6 Astra$10.00$50.00
GPT-5.6 Luna$0.20$1.20
GPT-5.6 Sol$4.00$20.00

GPT-6 Luna benchmark scores

GPT-6 Luna benchmark scores, provenance, and source links
BenchmarkScore
HLE(no tools)[1]
Reasoning · ±2 is noise
ARC-AGI-2(max)[2]
Compositional visual reasoning · ±9.2 is noise
LiveBench[3]
Composite score across 7 domains · ±2.7 is noise
DeepSWE[4]
Long-horizon coding · ±9.5 is noise

Who ran these numbers: 3 of 4 independent — artificialanalysis.ai (1), arcprize.org (1), livebench.ai (1); vendor self-reported (1).

  1. HLE: AA's own run of GPT-6 Luna at max effort, text-only 2,158-question subset.
  2. ARC-AGI-2: ARC Prize's own leaderboard row at the max effort tier ($0.062/task); v2.json dataset generated 2026-09-22.
  3. LiveBench: Board row 'GPT-6 Luna Max Effort' on the 2026-06-25 LiveBench release; only tier populated for this model.
  4. DeepSWE: OpenAI's own launch-post figure (max effort). No independent DeepSWE run for this model has appeared on deepswe.datacurve.ai or via Artificial Analysis's own agent-harness runs yet — unlike GPT-6 Sol, which AA has run through its Codex agent.

Notes on the record

Pricing is tiered by prompt length, confirmed on developers.openai.com/api/docs/models/gpt-6-luna: $0.10 input / $0.50 output per million tokens under 272,000 prompt tokens (cached input $0.01), rising to $0.20/$0.75 (cached $0.02) — 2x input and cache, 1.5x output, billed on the full request — once a prompt crosses that threshold, per OpenAI's own wording. OpenAI's comparison table reads "$0.20→$0.10 input, $1.20→$0.50 output, 50% cheaper" against GPT-5.6 Luna, but calls that $0.20/$1.20 rate "promotional pricing" — GPT-6 Luna's own price is presented as standard, not a further discount on a discount.

Context window 1,050,000 tokens, max output 128,000. Input: text and images. Output: text only. Knowledge cutoff May 18, 2026. License not stated for GPT-6 Luna; inferred proprietary, API-only.

No OpenAI page calls this a replacement for GPT-5.6 Luna. Beyond the absence of "supersede"/"replace" language, GPT-5.6 Luna is explicitly still in active use elsewhere on OpenAI's own site — its API deprecations page lists GPT-5.6 Luna as the recommended replacement for two older, separately-deprecating models.

OpenAI self-reports 66.6% on DeepSWE v1.1 (max effort); no independent DeepSWE run for this specific model has appeared yet on deepswe.datacurve.ai or Artificial Analysis's own agent-harness runs (unlike GPT-6 Sol, which AA has run through its Codex-agent harness).

Compare with

FAQ

Has GPT-6 Luna been independently benchmarked?

Partly, one day after launch — three of the eleven benchmarks this site tracks carry an independently-run score: Humanity's Last Exam (38.5%, Artificial Analysis, max effort), ARC-AGI-2 (59.31%, ARC Prize, max tier), and LiveBench (72.0 overall, max effort). GPQA Diamond, SWE-bench Verified, LiveCodeBench, Terminal-Bench 2.1, DeepSWE, Toolathlon-Verified, Agents' Last Exam and HMMT Feb 2026 have none yet — deepswe.datacurve.ai, tbench.ai, toolathlon.xyz, snorkel.ai, matharena.ai and vals.ai were all checked directly.

Is GPT-6 Luna better than GPT-5.6 Luna?

Not measurably. On all three benchmarks where both have independent scores from the same evaluator at max effort, GPT-6 Luna comes out slightly lower — Humanity's Last Exam 38.5% versus 39.5%, LiveBench 72.0 versus 73.6, ARC-AGI-2 59.31% versus 59.6% — and every gap sits inside that benchmark's noise band, so all three are ties. OpenAI's own improvement claim for Luna, 5.4 points on AutomationBench, is on a benchmark this site doesn't track. What clearly changed is price: $0.10/$0.50 versus $0.20/$1.20.

Does GPT-6 Luna's price change with longer prompts?

Yes. The $0.10 input / $0.50 output per-million-token rate applies to prompts up to 272,000 tokens. Once a prompt exceeds that threshold, OpenAI bills the whole request at 2x the input and cached-input rate and 1.5x the output rate — $0.20 input / $0.02 cached / $0.75 output — per developers.openai.com, confirmed 2026-09-23.

Does GPT-6 Luna replace GPT-5.6 Luna?

No. OpenAI's own API deprecations page lists GPT-5.6 Luna as the recommended replacement for two other, older models that are themselves being retired — meaning OpenAI is actively steering traffic toward GPT-5.6 Luna elsewhere on the same site, not away from it. Neither model's page uses "supersede" or "replace" to describe the relationship between them.

How does GPT-6 Luna compare to GPT-6 Sol?

Both launched the same day at the same tiered-pricing structure, but Luna is roughly a twentieth of Sol's price ($0.10/$0.50 versus $2/$10) and scores lower on every independent benchmark this page tracks for both — Humanity's Last Exam 38.5% versus 47.9%, LiveBench 72.0 versus 79.3 overall. OpenAI positions Luna as the cheaper, smaller sibling in the same GPT-6 family as Sol and Astra.

Further reading

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.