xAI (SpaceXAI)
Grok 4.6
All six scores tracked here are independent Vals.ai, Artificial Analysis, and DeepSWE-leaderboard runs, not xAI's own numbers — unusually clean for a model this new. What's still unverified is coverage, not accuracy: no tbench.ai, Toolathlon, or Snorkel score exists for it yet, and the advertised $2/$6 price is a sub-200k-token tier that doubles to $4/$12 above that threshold.
Grok 4.6’s 6 benchmark scores on this page were each verified against their sources on or after 2026-08-17.
- Released
- 2026-08-12
- License
- proprietary
- Context window
- 500K tokens
- Knowledge cutoff
- 2026-02
The verified record
Against the 76 head-to-head comparisons Grok 4.6 shares with other tracked models: 16 real gaps, 25 inside the noise band, and 35 we will not call.
A gap counts for Grok 4.6 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Grok 4.6 trails on 7 of them.
HLE · no toolsReasoning
±2 is noiseLiveCodeBenchContest coding
±3.1 is noiseDeepSWELong-horizon coding
±9.5 is noiseNo verdict for Grok 4.6 anywhere on Terminal-Bench 2.1 (nothing independently confirmed on both sides); GPQA Diamond, SWE-bench Verified (saturated).
- None of Grok 4.6’s agentic comparisons are independently confirmed on both sides yet.
Grok 4.6 API pricing
$2.00 in / $6.00 out per 1M tokens — official pricing
Grok 4.6 benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise | |
| SWE-bench Verifiedsaturated[2] Bug fixing — not ranked at any gap size | |
| GPQA Diamondsaturated[3] Expert science Q&A — not ranked at any gap size | |
| LiveCodeBench[4] Contest coding · ±3.1 is noise | |
| DeepSWE[5] Long-horizon coding · ±9.5 is noise | |
| Terminal-Bench 2.1[6] Terminal ops · ±10.6 is noise |
Who ran these numbers: 6 of 6 independent — vals.ai (3), artificialanalysis.ai (2), deepswe.datacurve.ai (1).
- HLE: AA's own run at HIGH effort (not xhigh) — text-only 2,158-question subset. Ties GPT-5.6 Terra.
- SWE-bench Verified: vals.ai run, rank 4/83+, bash-only harness, $0.78/test (updated 2026-08-14).
- GPQA Diamond: vals.ai run (94.697), rank 3 (updated 2026-08-15).
- LiveCodeBench: vals.ai run (88.224), updated 2026-08-15.
- DeepSWE: 67%±2 on mini-swe-agent (board updated 2026-08-13). xAI's own announcement claimed 65.9 — independent run came out slightly higher.
- Terminal-Bench 2.1: AA's own run of 'Grok 4.6 (high)' (0.883895131086142 in AA's payload). Not on tbench.ai's official board, which lists only the older Grok 4.5.
Notes on the record
Pricing shown is the sub-200k-prompt tier; above 200k tokens all rates double — see the FAQ below for the exact figures (confirmed on docs.x.ai/developers/models, verified 2026-08-20). Reasoning effort low/medium/high (default)/xhigh — confirmed current on xAI's reasoning docs (docs.x.ai/developers/model-capabilities/text/reasoning, verified 2026-08-20); xhigh is only available on Grok 4.6 and later — requests for it on Grok 4.5 fall back to high.
Supersedes Grok 4.5 (July 2026). Announcement verified at x.ai/news/grok-4-6. xAI's own news page now brands itself "SpaceXAI": SpaceX acquired xAI on 2026-02-02 (source: x.ai/news/xai-joins-spacex, widely reported by Reuters/Bloomberg/CNBC), and xAI's channels formally adopted the "SpaceXAI" name on 2026-07-06 — both predate Grok 4.6, and Grok 4.5 (July 2026) already shipped under this same combined corporate structure, so the "xAI (SpaceXAI)" lab label reflects the company's current name, not anything new about this release.
Not yet on tbench.ai, Toolathlon, or Snorkel boards — all predate its release; as of 2026-08-20 (8 days post-launch) none of the three had published a Grok 4.6 score.
Compare with
FAQ
Has Grok 4.6 been independently benchmarked?
Yes, on all six metrics Grok 4.6 has scores for on this site. HLE, Terminal-Bench 2.1, SWE-bench Verified, GPQA Diamond, LiveCodeBench, and DeepSWE are each independent runs — from Vals.ai, Artificial Analysis, and the DeepSWE leaderboard (DataCurve), observed 2026-08-17 and 2026-08-20 — not figures self-reported by xAI. That's 6 of 6 independently sourced; see the table above for the individual figures.
Why is there no tbench.ai, Toolathlon, or Snorkel score for Grok 4.6?
Grok 4.6 released 2026-08-12, and none of those three agentic-benchmark boards had published a run for it as of 2026-08-20. All three boards predate the model's release, so the gap reflects timing, not a refused or failed evaluation.
Does Grok 4.6's price change with longer prompts?
Yes. The $2 input / $6 output per-million-token rate shown on this page applies only to prompts under 200,000 tokens. Above that threshold, xAI bills the entire request at the higher tier — $4 input / $12 output per million tokens, per docs.x.ai/developers/models (verified 2026-08-20). Cached input pricing doubles the same way, from $0.50 to $1.00 per million tokens, at the same 200k cutoff.
What reasoning effort levels does Grok 4.6 support?
Four settings — low, medium, high (the default), and xhigh — confirmed on xAI's reasoning docs (docs.x.ai/developers/model-capabilities/text/reasoning, verified 2026-08-20); xhigh is only available on Grok 4.6 and later, so requests for it on Grok 4.5 fall back to high. Higher settings trade latency and token spend for more deliberate reasoning; the benchmark scores on this page reflect whichever effort tier the source evaluation used, not necessarily the default.
Why does this page list Grok 4.6's lab as "xAI (SpaceXAI)"?
SpaceX acquired xAI in an all-stock merger on 2026-02-02, and xAI's own channels rebranded to "SpaceXAI" on 2026-07-06 (source: x.ai/news/xai-joins-spacex, widely reported by Reuters/Bloomberg/CNBC). Both changes predate Grok 4.6: Grok 4.5 (July 2026) was already released under the same combined corporate structure, so Grok 4.6 is not the first model shipped this way — the "xAI (SpaceXAI)" label here just reflects the lab's current corporate name, not anything new about this specific release.