OpenAI

GPT-5.6 Luna

All eight scores tracked here for GPT-5.6 Luna come from independent evaluators, not OpenAI's own numbers — trustworthy on sourcing, but not fully consistent: the same Terminal-Bench 2.1 test reads 80.9% on Artificial Analysis's harness versus 79.03% on vals.ai's own Terminus-2 run of the identical model, and LiveCodeBench has no Luna entry on vals.ai's tracker, and Toolathlon-Verified and HMMT Feb 2026 have no Luna row on their own official leaderboards.

GPT-5.6 Luna’s 8 benchmark scores on this page were verified against their sources on or after 2026-08-26.

Released
2026-07-09
License
proprietary
Context window
1M tokens
Knowledge cutoff
2026-02-16

The verified record

Against the 107 head-to-head comparisons GPT-5.6 Luna shares with other tracked models: 36 real gaps, 28 inside the noise band, and 43 we will not call.

A gap counts for GPT-5.6 Luna only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — GPT-5.6 Luna trails on 24 of them.

LiveBenchComposite score across 7 domains

±2.7 is noise
Ahead
Behind
Claude Fable 5 9.4 · GPT-5.6 Sol 7.4 · Claude Opus 5 6.5 · Kimi K3 5.6 · Gemini 3.7 Flash 5.2 · Qwen3.8-Max 4.9 · Grok 4.6 4.4 · Muse Spark 1.2 4.4 · DeepSeek V4 Pro (0813) 3.8 · Gemini 3.1 Pro Preview 3.4
Tie
5 models within ±2.7

HLE · no toolsReasoning

±2 is noise
Behind
Claude Fable 5 16.0 · Claude Opus 5 15.4 · GPT-5.6 Sol 10.0 · Claude Opus 4.8 9.2 · Gemini 3.7 Flash 8.4 · Gemini 3.1 Pro Preview 7.5 · Kimi K3 7.4 · Muse Spark 1.2 6.0 · Qwen3.8-Max 3.5 · Grok 4.6 3.4 · GLM-5.3 2.8
Tie
4 models within ±2

Agents' Last ExamProfessional work

±3.2 is noise
Ahead
Tie
2 models within ±3.2
Unverified
1 model — vendor-reported on one side

DeepSWELong-horizon coding

±9.5 is noise
Ahead
Tie
7 models within ±9.5
Unverified
2 models — vendor-reported on one side

ARC-AGI-2 · maxCompositional visual reasoning

±9.2 is noise
Behind
GPT-5.6 Sol 32.9 · Claude Opus 5 30.8 · Claude Fable 5 29.6
Tie
3 models within ±9.2

No verdict for GPT-5.6 Luna anywhere on Terminal-Bench 2.1 (nothing independently confirmed on both sides); GPQA Diamond, SWE-bench Verified (saturated).

GPT-5.6 Luna API pricing

$0.20 in / $1.20 out per 1M tokens official pricing

GPT-5.6 Luna is one of 2 OpenAI models tracked on this site, at these official list prices.

OpenAI model pricing, official list rates
ModelIn / 1MOut / 1M
GPT-5.6 Luna$0.20$1.20
GPT-5.6 Sol$5.00$30.00

GPT-5.6 Luna benchmark scores

GPT-5.6 Luna benchmark scores, provenance, and source links
BenchmarkScore
GPQA Diamondsaturated[1]
Expert science Q&A — not ranked at any gap size
SWE-bench Verifiedsaturated[2]
Bug fixing — not ranked at any gap size
HLE(no tools)[3]
Reasoning · ±2 is noise
Terminal-Bench 2.1[4]
Terminal ops · ±10.6 is noise
DeepSWE[5]
Long-horizon coding · ±9.5 is noise
Agents' Last Exam[6]
Professional work · ±3.2 is noise
ARC-AGI-2(max)[7]
Compositional visual reasoning · ±9.2 is noise
LiveBench[8]
Composite score across 7 domains · ±2.7 is noise

Who ran these numbers: 8 of 8 independent — vals.ai (2), artificialanalysis.ai (2), deepswe.datacurve.ai (1), snorkel.ai (1), arcprize.org (1), livebench.ai (1).

  1. GPQA Diamond: Rank 16/135 on vals.ai; benchmark flagged "largely saturated" by vals.ai (24/135 models ≥90%). Cost $0.2/$1.2, latency 53.52s.
  2. SWE-bench Verified: Rank 9/86 on vals.ai's minimal bash-tool-only harness. Cost $0.04/test, latency 3m21s.
  3. HLE: Reported at the model's "max" reasoning-effort setting. Rank ~15 of 31 models shown. AA's chart renders no plain-text data table; value was read by matching the model's x-axis tick position to the corresponding bar-value label (rank-order + pixel-offset cross-check against neighboring known models).
  4. Terminal-Bench 2.1: Reported at the model's "max" reasoning-effort setting. AA runs Terminus 2 harness in an e2b sandbox, pass@1 averaged over 3 repeats. Same value-extraction method as HLE. Cross-check: vals.ai's own Terminus-2 run of the same model gives 79.03%±0.99 (not used here since vals-ai isn't an approved source for this benchmark id).
  5. DeepSWE: Reported at the model's max-effort tier, DeepSWE v1.1 Best-of-effort scoring. 67%±4%. Avg cost $0.61/task, 73k output tokens, 102 agent steps. 113-task corpus, leaderboard updated 2026-08-20. All models run on mini-swe-agent via Pier.
  6. Agents' Last Exam: Reported using the Codex harness at the model's "XHigh" effort tier (best-of-effort scoring). Value is Pass Rate (share of runs with a perfect score); partial-credit Score for the same run is 49.4%. Est. cost $235 over 66h7m total runtime. "Best-per-task" snapshot dated 2026-07-04.
  7. ARC-AGI-2: Cost/task $0.177. Same entry's ARC-AGI-1 score is 90.7%. Entry is dated 2026-07-30 (three weeks post-launch, same date as Luna's 80% price cut) rather than the 2026-07-09 release date.
  8. LiveBench: Reported at the model's "Max Effort" setting. Overall score on LiveBench-2026-06-25 release. Sub-scores: Reasoning 85.6, Coding 82.9, Agentic Coding 48.4, Math 87.2, Data Analysis 78.0, Language 72.6, Instruction Following 60.1. Cost $0.169/successful task.

Notes on the record

GPT-5.6 Luna launched 2026-07-09 alongside two siblings, Sol and Terra, as the cheapest and fastest tier of the family — OpenAI's rough "nano" analogue for this generation, positioned for high-volume, latency-sensitive workloads rather than peak capability. It carries a 1,050,000-token context window and a 2026-02-16 knowledge cutoff. All three siblings spent from 2026-06-26 in a restricted preview limited to roughly 20 US-government-vetted organizations while the Commerce Department's AI standards center completed a review; general availability followed once that review closed.

Pricing has three distinct tiers, not one number. Standard short-context is $0.20 in / $1.20 out per 1M tokens, with cached input at $0.02; long-context requests above 272K input tokens run at 2x input and 1.5x output ($0.40/$1.80); batch/flex asynchronous processing runs at 50% of standard ($0.10/$0.60) (OpenAI API pricing docs). Multiple secondary trackers report OpenAI cut Luna's price by roughly 80% on 2026-07-30, three weeks after launch — that pre-cut figure was not captured live by this site, so it is flagged rather than stated as verified.

All eight benchmark scores logged for Luna here come from independent evaluators rather than OpenAI, but coverage is uneven and at least two numbers need a caveat before they're read at face value. GPQA Diamond's 91.67% (vals.ai, rank 16/135, as of 2026-08-26) sits on a benchmark vals.ai itself flags as "largely saturated" — 24 of 135 tracked models already clear 90%. Humanity's Last Exam (39.5%, "max" variant) and Terminal-Bench 2.1 (80.9%, "max" variant) both come from Artificial Analysis, whose chart renders no plain-text data table; both values were read via pixel-offset/rank-order matching against neighboring known models rather than a direct number lookup. Terminal-Bench 2.1 also has a second, lower reading: vals.ai's own Terminus-2 run of the identical model gives 79.03%±0.99, a roughly 1.9-point gap from AA's figure attributable to implementation/sampling differences between the two independent harnesses — this site's benchmark-id rule uses only the AA figure for the terminal-bench-2-1 field. Agents' Last Exam's 30.3% is a Pass Rate (share of runs scoring perfectly); the same Codex/XHigh run's partial-credit Score is materially higher at 49.4%, so the two should not be read as interchangeable. The ARC-AGI-2 entry (59.6%; the same run's ARC-AGI-1 is 90.7%) is dated 2026-07-30 on the ARC Prize leaderboard — the same date as the reported price cut, not the 2026-07-09 release date, so it may reflect a later checkpoint rather than the model as it originally shipped. LiveBench's Max Effort run scores 73.6% Overall (Reasoning 85.6, Coding 82.9, Agentic Coding 48.4, Mathematics 87.2, Data Analysis 78.0, Language 72.6, Instruction Following 60.1; $0.169/successful task, LiveBench-2026-06-25 release, as of 2026-08-26). Three benchmarks tracked elsewhere on this site were checked and found to have no Luna row at all, rather than a low or estimated one: LiveCodeBench (absent from vals.ai's Luna page), Toolathlon-Verified (absent for the entire GPT-5.6 family on its official leaderboard), and MathArena's HMMT Feb 2026 table (no GPT-5.6 entries for any tier, as of 2026-08-26).

Luna's clearest signal in the data is cost efficiency relative to its own family, not raw ranking. It clears SWE-bench Verified at 93.00% for $0.04/test versus Sol's reported 96.20% for $1.15/test — about 29x cheaper for a 3.2-point accuracy gap. On DeepSWE v1.1 it scores 67%±4% for $0.61/task versus Sol's 73% for $6.46/task, about 11x cheaper for a 6-point gap. On Agents' Last Exam its Codex/XHigh run totals $235 for a 30.3% pass rate versus Sol's $772 for 30.6%, about 3x cheaper for a marginal gap (vals.ai, DeepSWE, and Agents' Last Exam leaderboards, as of 2026-08-26). Licensing is proprietary: weights are marked "Private" on vals.ai and "Proprietary" on Artificial Analysis.

Compare with

FAQ

Is GPT-5.6 Luna good for coding?

On the two independent coding benchmarks tracked here, yes for the price: SWE-bench Verified 93.00% (rank 9/86 on vals.ai's minimal bash-tool-only harness, $0.04/test) and DeepSWE v1.1 67%±4% ($0.61/task, 113-task corpus, updated 2026-08-20). Both trail its costlier sibling Sol (96.20% and 73% respectively) by a few points, but Sol costs roughly 10-29x more per task on these two benchmarks, so Luna reads as the better fit for high-volume coding workloads rather than the highest achievable score.

How much does GPT-5.6 Luna cost to use?

Standard API pricing is $0.20 per 1M input tokens and $1.20 per 1M output tokens, with cached input at $0.02. Two other tiers apply: long-context requests above 272K input tokens cost 2x input / 1.5x output ($0.40/$1.80), and batch/flex asynchronous processing runs at half the standard rate ($0.10/$0.60). Secondary trackers report an additional roughly 80% price cut on 2026-07-30, three weeks after launch, though this site did not capture the pre-cut price live to verify it independently.

Why do Terminal-Bench 2.1 scores for GPT-5.6 Luna differ between sources?

Artificial Analysis's independent run (e2b sandbox, pass@1 averaged over 3 repeats) puts Luna at 80.9%, while vals.ai's own run of the same Terminus-2 agent against the same model gives 79.03%±0.99. Both are legitimate independent measurements of the identical model and agent family; the roughly 1.9-point gap comes from implementation and sampling differences between the two evaluators' setups, not from a different model version. This site reports the Artificial Analysis figure in its main terminal-bench-2-1 field.

Has GPT-5.6 Luna been benchmarked by OpenAI itself, or only by outside evaluators?

Every one of the eight benchmark scores logged here for Luna comes from an outside evaluator rather than an OpenAI-published number, spanning reasoning, coding, agentic, and abstract-reasoning tests. That said, coverage isn't complete: LiveCodeBench has no Luna entry on vals.ai's tracker (its own official leaderboard wasn't separately checked), and Toolathlon-Verified and MathArena's HMMT Feb 2026 have no Luna row on their own official leaderboards — so those gaps are left blank rather than guessed at.

How does GPT-5.6 Luna compare to its sibling GPT-5.6 Sol?

Luna is the cheapest, fastest tier of the GPT-5.6 family, and the numbers back that positioning: it trails Sol by only a few accuracy points on shared benchmarks (e.g., 93.00% vs 96.20% on SWE-bench Verified) while costing dramatically less per run — about 29x less per SWE-bench test ($0.04 vs $1.15) and roughly 3x less per Agents' Last Exam run ($235 vs $772). For high-volume or latency-sensitive use, that trade generally favors Luna over Sol.

Why is GPT-5.6 Luna's ARC-AGI-2 score dated three weeks after its release?

The ARC Prize leaderboard lists Luna's ARC-AGI-2 entry (59.6%, $0.177/task; the same entry scores 90.7% on ARC-AGI-1) under the date 2026-07-30, not the model's 2026-07-09 launch date. That date happens to match the roughly 80% price cut reported by secondary trackers around the same time, which suggests the scored run may reflect a later checkpoint or configuration — though this site cannot confirm that link beyond the date coincidence.

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.