OpenAI
GPT-6 Astra
OpenAI calls GPT-6 Astra its most capable model yet and the first to cross its own "Critical" cybersecurity threshold — but of the five tracked benchmarks with an independent score so far, only Humanity's Last Exam separates it from GPT-5.6 Sol on a clean same-tier, same-harness comparison, by 5.2 points; the only other clean comparison, ARC-AGI-2, lands inside that benchmark's noise band. Six of the eleven boards this site tracks hadn't scored it at all as of the day after launch.
GPT-6 Astra benchmarks and pricing, every number sourced: 5 tracked GPT-6 Astra benchmark scores (5 independently run, 0 still resting on a vendor’s own claim), priced at $10.00 per million input tokens and $50.00 per million output.
GPT-6 Astra’s 5 benchmark scores on this page were verified against their source on 2026-09-04.
- Released
- 2026-09-03
- License
- proprietary
- Context window
- 1M tokens
- Knowledge cutoff
- 2026-04-30
- Parameters
- Not disclosed
- Architecture
- Not disclosed
GPT-6 Astra’s verified record
Against the 90 head-to-head comparisons GPT-6 Astra shares with other tracked models: 32 real gaps, 17 inside the noise band, and 41 we will not call.
A gap counts for GPT-6 Astra only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — GPT-6 Astra trails on 1 of them.
HLE · no tools Reasoning
±2 is noiseDeepSWE Long-horizon coding
±9.5 is noiseARC-AGI-2 · max Compositional visual reasoning
±9.2 is noiseTerminal-Bench 2.1 Terminal ops
±10.6 is noiseNo verdict for GPT-6 Astra anywhere on GPQA Diamond (saturated).
GPT-6 Astra API pricing
$10.00 in / $50.00 out per 1M tokens — official pricing
What GPT-6 Astra costs per job
| Workload | Tokens in / out | Cost |
|---|---|---|
| One long chat turn | 100K / 10K | $1.50 |
| A codebase review | 1,000K / 100K | $15.00 |
| A day of agent work | 10,000K / 1,000K | $150.00 |
Computed from GPT-6 Astra’s list rates above — cache discounts and batch tiers are not applied.
GPT-6 Astra is one of 3 OpenAI models tracked on this site, at these official list prices.
| Model | In / 1M | Out / 1M |
|---|---|---|
| GPT-6 Astra | $10.00 | $50.00 |
| GPT-5.6 Luna | $0.20 | $1.20 |
| GPT-5.6 Sol | $4.00 | $20.00 |
GPT-6 Astra benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise | |
| GPQA Diamondsaturated[2] Expert science Q&A — not ranked at any gap size | |
| Terminal-Bench 2.1[3] Terminal ops · ±10.6 is noise | |
| DeepSWE[4] Long-horizon coding · ±9.5 is noise | |
| ARC-AGI-2(max)[5] Compositional visual reasoning · ±9.2 is noise |
Who ran these numbers: 5 of 5 independent — artificialanalysis.ai (2), tbench.ai (1), deepswe.datacurve.ai (1), arcprize.org (1).
- HLE: AA's own run at max effort, text-only subset (raw payload value 0.546802594995366). AA also lists xhigh (54.59%), high, medium, low, and non-reasoning tiers for this model.
- GPQA Diamond: AA's own run at max effort (0.960606060606061 in AA's payload; independently confirmed via a direct curl of the same URL). AA also lists xhigh (96.26% — its single highest GPQA Diamond score to date, a tier-naming inversion), high (94.95%), and medium (93.94%). This benchmark is saturated on this site; every comparison on it is tainted regardless of gap size.
- Terminal-Bench 2.1: tbench.ai's own canonical board, "GPT-6 Astra (high)" run via the Codex agent, rank 1 of 18 (±1.8%, 95% CI). Artificial Analysis has not added a Terminal-Bench 2.1 score for this model — its own GPT-6 Astra page instead headlines a separate benchmark, "Terminal-Bench 4.0" (57.9%, max tier), which this site does not track. GPT-5.6 Sol's own tracked score for this benchmark comes from Artificial Analysis, not tbench.ai — a different harness, recorded as such.
- DeepSWE: 74%±3% at xhigh effort (the model's own best config on this board's v1.1 table, $6.52/task, 30k output tokens, 29 steps), tied for the board's top spot with Gemini 3.8 Flash (high) and Claude Opus 5 (max). GPT-5.6 Sol's own best config on the same board is max effort, 73%±3% — a cross-tier comparison by construction, though checking Astra's own max-tier row (73%±1%, per the board's full effort-level view) shows the two are functionally tied at matched tiers too.
- ARC-AGI-2: GPT-6 Astra's official ARC-AGI-2 leaderboard row (Max tier), 95.0% — 2.5 points above GPT-5.6 Sol's own tracked 92.5% at the same tier, a gap this benchmark's 9.2-point noise floor swallows as a tie.
Notes on the record
GPT-6 Astra launched September 3, 2026, per OpenAI's own announcement blog, safety overview, and System Card — all three independently corroborated by Bloomberg, CNBC, 9to5Mac, and Axios coverage the same day. OpenAI frames it as "the world's most intelligent and aligned model" and "the most capable model we have ever broadly deployed," positioned above the existing GPT-5.6 line (GPT-5.6 Sol, Terra, and Luna) rather than replacing it: none of OpenAI's own model or deprecation pages describes Astra as a successor to any GPT-5.6 model, and all three remain sold at unchanged prices with no retirement date — unlike this site's established supersession cases (Claude Fable 5 to Fable 5.1, Gemini 3.7 Flash to 3.8 Flash), where the vendor's own comparison table used explicit "successor to" language. Astra ships as a single model with five reasoning-effort settings (low/medium/high/xhigh/max) rather than as a three-tier family the way GPT-5.6 did. OpenAI's own launch page reinforces the same positioning without using supersession language: its benchmark tables compare Astra exclusively against GPT-5.6 Sol, and its own "choosing a model" guidance now recommends Astra first and omits Sol from that sentence entirely (while still selling it) — informal signals Astra has taken the flagship slot, not a formal retirement.
API pricing is $10 per million input tokens and $50 per million output tokens for prompts up to 272,000 tokens, per OpenAI's own pricing page — the same 272K threshold GPT-5.6 Sol and Luna already carry on this site. Above that line, the entire request bills at double the input rate and 1.5 times the output rate ($20/$75), a rule OpenAI states applies to the full request, not just the tokens past the line. Context window is 1,050,000 tokens, matching Sol and Luna's own figure exactly; max output is 128,000 tokens. Knowledge cutoff is April 30, 2026, both stated in plain text on OpenAI's own model page. No parameter count is disclosed, consistent with every other OpenAI model tracked here.
One day after launch, five of the eleven benchmarks this site tracks carry an independently-run Astra score: Humanity's Last Exam (54.68% at max effort, Artificial Analysis), GPQA Diamond (96.06% at max effort, also Artificial Analysis — this benchmark is saturated on this site and every comparison on it is tainted regardless of gap size), Terminal-Bench 2.1 (87.4% at high effort, tbench.ai's own canonical board, rank 1 of 18), DeepSWE (74% at xhigh effort, tied for the board's top spot), and ARC-AGI-2 (95.0% at Max tier, arcprize.org). Six others — LiveCodeBench, SWE-bench Verified, Toolathlon-Verified, Agents' Last Exam, LiveBench, and HMMT Feb 2026 — had not added Astra as of 2026-09-04, checked live on each; all six boards' most recent update predates the model's release.
Only one of those five produces a clean, same-tier, same-harness comparison against a tracked GPT-5.6 model: Humanity's Last Exam, where Astra's 54.68% clears GPT-5.6 Sol's own 49.49% (both at max effort, both Artificial Analysis) by 5.19 points — a real gap against this benchmark's 2-point noise floor. ARC-AGI-2 is also same-tier and same-harness (both Max, both arcprize.org), but the 2.5-point gap over Sol's 92.5% sits inside that benchmark's 9.2-point margin — a tie, not the confident jump the raw scores alone would suggest. DeepSWE compares Astra's own best config (74% at xhigh) against Sol's own best (73% at max) — a mismatched-tier comparison by construction, though checking Astra at the same max tier Sol uses (73%, per deepswe.datacurve.ai's full effort-level table) shows the two are functionally tied regardless of which tier is used. Terminal-Bench 2.1 can't be compared at all: Astra's only tracked score comes from tbench.ai's canonical board, while Sol's comes from Artificial Analysis's own run — two different harnesses on the same benchmark, which this site records as Setup-dependent rather than guessing which one would win on the other's turf.
OpenAI's own System Card discloses that Astra is the first model to cross the "Critical" cybersecurity capability threshold under its Preparedness Framework — confirmed on both the pre-launch "Path to Astra" post (2026-09-01) and the System Card itself (2026-09-03), not just press paraphrase. The evidence OpenAI cites: 100% on its internal ExploitBench (versus 78.5% for Sol), including two previously-unknown zero-days found on a version of that benchmark built entirely from post-cutoff CVEs to rule out contamination; and internal red-team exercises where Astra, given only a starting goal, built a full sandbox-escape exploit chain against a hardened browser and a privilege-escalation chain against a hardened OS kernel — both previously undisclosed vulnerabilities, both since reported to maintainers. Disclosed mitigations include a refusal boundary tuned to block proof-of-concept exploit requests by default (91.5% refusal rate on an internal cyber-jailbreak set, versus 59% for Sol) and a production monitoring layer watching tool-using sessions across ChatGPT, Codex, and the API. OpenAI also discloses, without spinning it as solved, that Astra's own monitorability has gone down relative to Sol: in adversarial tests explicitly instructing it to evade detection, it can sometimes do so undetected on sabotage-style tasks.
A more capable, safety-loosened sibling exists for vetted cyber-defense customers, but it is not Astra: OpenAI's "Daybreak Red" access tier — the one that unlocks proof-of-concept exploit generation — runs on a separate, earlier model called GPT-5.6-Cyber, not on Astra. A lighter tier for Astra itself, "Daybreak Blue," is not yet broadly available either: OpenAI's own Help Center states reduced refusals aren't available on Astra for most Daybreak customers, and the benchmark figures OpenAI cites for Astra under Daybreak Blue (a 92% proof-of-concept-exploit completion rate, versus 2.4% under standard safeguards) came from a special internal test configuration OpenAI itself labels not the default production setup — not from general access. OpenAI says it plans to expand Daybreak access for Astra in the coming weeks, without committing to a specific Astra-based Daybreak Red tier. Where this site has recorded a similar-looking restricted sibling for other vendors (Claude Mythos 5.1, Gemini 3.8 Flash Cyber), those are the same-generation model with fewer guardrails; Astra's cyber-capable counterpart today is a different, older model entirely, and Astra's own reduced-safeguard access is still rolling out.
Astra does not appear under its own name inside ChatGPT's model picker — it ships there as "GPT-6 Pro," a separate option from "GPT-5.6 Sol Pro" with its own usage caps that vary by plan, per OpenAI's own Help Center article. Two of OpenAI's own pages disagree on what to call the initial enterprise rollout program ("Trusted Access Program" on one, "Daybreak Access Program" on another) — a naming inconsistency this site is disclosing rather than resolving, since OpenAI itself hasn't. The same two pages also disagree on ChatGPT Plus: the announcement says Astra reaches all Plus users "over the coming days," while the Help Center states GPT-6 Pro is not included with Plus in Chat — reported here as the conflict it is.
Compare with
FAQ
Is GPT-6 Astra actually better than GPT-5.6 Sol?
On the one benchmark where a clean comparison actually separates them, yes: Humanity's Last Exam, where Astra's 54.68% beats Sol's 49.49% by 5.19 points at the same reasoning-effort tier and the same evaluator (Artificial Analysis) — a real gap against that benchmark's 2-point noise floor. The only other clean, matched-tier, same-evaluator comparison is ARC-AGI-2, where Astra leads by 2.5 points — inside the benchmark's own 9.2-point margin, so it counts as a tie. Everywhere else the comparison is muddier. DeepSWE and Terminal-Bench 2.1 can't be compared cleanly — DeepSWE only at mismatched effort tiers (though the scores end up tied even at matched tiers), and Terminal-Bench 2.1 only via two different evaluators (tbench.ai for Astra, Artificial Analysis for Sol). GPQA Diamond favors Astra too, but this site has stopped ranking that benchmark entirely — it's saturated, so the gap doesn't count regardless of size.
Why did OpenAI restrict GPT-6 Astra's cybersecurity capabilities?
Because OpenAI's own System Card says Astra is the first model to cross the "Critical" threshold for cybersecurity under its Preparedness Framework — evidence includes a 100% score on OpenAI's internal ExploitBench (versus 78.5% for GPT-5.6 Sol), including two previously-unknown zero-day vulnerabilities the model found and that have since been reported to maintainers. By default, Astra refuses advanced cybersecurity tasks like building proof-of-concept exploits (a 91.5% refusal rate on OpenAI's own cyber-jailbreak eval, versus 59% for Sol), and a production monitoring layer watches tool-using sessions in ChatGPT, Codex, and the API. OpenAI also discloses, without downplaying it, that Astra's own monitorability — how reliably its reasoning can be watched for deception — has gone down compared to Sol.
Does GPT-6 Astra replace GPT-5.6 Sol?
No — not according to OpenAI's own materials. None of OpenAI's model, pricing, or deprecation pages calls Astra a successor to Sol, Terra, or Luna; all three remain sold at unchanged prices with no retirement date. That's a different pattern from other supersessions this site has recorded, like Claude Fable 5 to Fable 5.1, where the vendor's own comparison table used explicit "successor to" language. OpenAI's own launch page does exclusively benchmark Astra against Sol, and its "choosing a model" guidance now recommends Astra first and omits Sol from that list entirely (while still selling it at full price) — informal signals that Astra has taken the flagship slot — but this site only marks a model superseded on explicit vendor language, and that language doesn't exist here.
How much does GPT-6 Astra cost, and does the price change for long prompts?
$10 per million input tokens and $50 per million output tokens for prompts up to 272,000 tokens — the same threshold GPT-5.6 Sol and Luna already carry on this site. Cross that line and the entire request, not just the excess, bills at double the input rate and 1.5 times the output rate ($20/$75 per million tokens), per OpenAI's own pricing page.
Has GPT-6 Astra been independently benchmarked?
Partially so far — five of the eleven benchmarks this site tracks carry an independently-run score one day after the model's 2026-09-03 launch: Humanity's Last Exam, GPQA Diamond, Terminal-Bench 2.1, DeepSWE, and ARC-AGI-2. LiveCodeBench, SWE-bench Verified, Toolathlon-Verified, Agents' Last Exam, LiveBench, and HMMT Feb 2026 had not added the model as of 2026-09-04, checked live on each. None of the five tracked scores here come from OpenAI's own claimed numbers.
Is GPT-6 Pro the same as GPT-6 Astra?
Yes — same underlying model, different name. In ChatGPT's model picker Astra does not appear under its own name; OpenAI's Help Center lists it as "GPT-6 Pro," available on the Pro ($100 and $200), Business, and Enterprise plans, with its own usage caps kept separate from "GPT-5.6 Sol Pro." The API name is gpt-6-astra. One caveat this site can't resolve: OpenAI's own announcement says Astra will reach all ChatGPT Plus users "over the coming days," while its Help Center states GPT-6 Pro is not included with Plus in Chat — two OpenAI pages, published the same window, disagreeing on Plus access.
Further reading
- Models with 10M token context windows 2026 — GPT-6 Astra is one of the 23 models it compares.