xAI (SpaceXAI)Previous version · Grok 4.7
Grok 4.6
All nine scores tracked here are independent Vals.ai, Artificial Analysis, DeepSWE-leaderboard, ARC Prize, and LiveBench runs, not xAI's own numbers — unusually clean for a model this new. What's still unverified is coverage, not accuracy: no tbench.ai, Toolathlon, or Snorkel score exists for it yet, and the advertised $2/$6 price is a sub-200k-token tier that doubles to $4/$12 above that threshold.
Grok 4.6 benchmarks and pricing, every number sourced: 9 tracked Grok 4.6 benchmark scores (9 independently run, 0 still resting on a vendor’s own claim), priced at $2.00 per million input tokens and $6.00 per million output.
- Released
- 2026-08-12
- License
- proprietary
- Context window
- 500K tokens
- Knowledge cutoff
- 2026-02
- Verified
- sources checked 2026-08-17–2026-09-29
- Parameters
- Not disclosed
- Architecture
- Not disclosed
Grok 4.6’s verified record
Grok 4.6’s most-compared rival is Kimi K3: 1 trail, 2 ties, and 5 not callable across their 8 shared comparisons. Grok 4.6 is priced at $2.00/$6.00 per 1M tokens (in/out) vs Kimi K3’s $3.00/$15.00.
Against the 195 head-to-head comparisons Grok 4.6 shares with other tracked models: 31 real gaps, 36 inside the noise band, and 128 we will not call.
A gap counts for Grok 4.6 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Grok 4.6 trails on 16 of them.
LiveBench Composite score across 7 domains
±2.7 is noiseHLE · no tools Reasoning
±2 is noiseDeepSWE Long-horizon coding
±9.5 is noiseAnalystAgent Spreadsheet & document analysis
±11.2 is noiseNo verdict for Grok 4.6 anywhere on Terminal-Bench 2.1 ( every independently confirmed comparison inside the noise band); ARC-AGI-2 · xhigh ( every comparison inside the noise band); GPQA Diamond, LiveCodeBench, SWE-bench Verified ( saturated).
Grok 4.6 API pricing
$2.00 in / $6.00 out per 1M tokens — official pricing
What Grok 4.6 costs per job
| Workload | Tokens in / out | Cost |
|---|---|---|
| One long chat turn | 100K / 10K | $0.260 |
| A codebase review | 1,000K / 100K | $2.60 |
| A day of agent work | 10,000K / 1,000K | $26.00 |
Computed from Grok 4.6’s list rates above — cache discounts and batch tiers are not applied.
Grok 4.6 is one of 2 xAI (SpaceXAI) models tracked on this site, at these official list prices.
| Model | In / 1M | Out / 1M |
|---|---|---|
| Grok 4.7 | $2.00 | $6.00 |
| Grok 4.6(previous version) | $2.00 | $6.00 |
Grok 4.6 benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise | |
| SWE-bench Verifiedsaturated[2] Bug fixing — not ranked at any gap size | |
| GPQA Diamondsaturated[3] Expert science Q&A — not ranked at any gap size | |
| LiveCodeBenchsaturated[4] Contest coding — not ranked at any gap size | |
| DeepSWE[5] Long-horizon coding · ±9.5 is noise | |
| Terminal-Bench 2.1[6] Terminal ops · ±10.6 is noise | |
| ARC-AGI-2(xhigh)[7] Compositional visual reasoning · ±9.2 is noise | |
| LiveBench[8] Composite score across 7 domains · ±2.7 is noise | |
| AnalystAgent[9] Spreadsheet & document analysis · ±11.2 is noise |
Who ran these numbers: 9 of 9 independent — artificialanalysis.ai (3), vals.ai (3), deepswe.datacurve.ai (1), arcprize.org (1), livebench.ai (1).
- HLE: AA's own run of Grok 4.6 at HIGH effort (not xhigh) — text-only 2,158-question subset. Ties GPT-5.6 Terra.
- SWE-bench Verified: vals.ai run, rank 4/83+, bash-only harness, $0.78/test (updated 2026-08-14).
- GPQA Diamond: vals.ai run (94.697), rank 3 (updated 2026-08-15).
- LiveCodeBench: vals.ai run (88.224), updated 2026-08-15. Corrected 2026-10-03: the value field read 88.2 (truncated); the board displays 88.22.
- DeepSWE: 67%±2 on mini-swe-agent (board updated 2026-08-13). xAI's own announcement claimed 65.9 — independent run came out slightly higher.
- Terminal-Bench 2.1: AA's own run of 'Grok 4.6 (high)'. Not on tbench.ai's official board, which lists only the older Grok 4.5. Cross-check (2026-10-01): vals.ai's archived Terminus 2 table lists it at 78.28.
- ARC-AGI-2: Grok 4.6's official ARC-AGI-2 leaderboard row, dated 2026-08-11 on arcprize.org. No Max tier exists for this model on the board — XHigh is its ceiling, the top of four populated tiers.
- LiveBench: Board row "Grok 4.6" on the 2026-06-25 LiveBench release.
- AnalystAgent: AA's own run, board row 'Grok 4.6 (high)' — same high tier as this site's AA HLE row for this model. Grok 4.6's value was read from AA's model page on 2026-09-29; the leaderboard's embedded payload carries only its 14 default-selected models, which excludes this one.
Notes on the record
Pricing shown is the sub-200k-prompt tier; at 200k tokens and above (xAI's own row label reads ">= 200k", not just ">") all rates double — see the FAQ below for the exact figures (confirmed on docs.x.ai/developers/models, verified 2026-08-20). Reasoning effort low/medium/high (default)/xhigh — confirmed current on xAI's reasoning docs (docs.x.ai/developers/model-capabilities/text/reasoning, verified 2026-08-20); xhigh is only available on Grok 4.6 and later — requests for it on Grok 4.5 fall back to high.
Supersedes Grok 4.5 (July 2026). Announcement verified at x.ai/news/grok-4-6. xAI's own news page now brands itself "SpaceXAI": SpaceX acquired xAI on 2026-02-02 (source: x.ai/news/xai-joins-spacex, widely reported by Reuters/Bloomberg/CNBC), and xAI's channels formally adopted the "SpaceXAI" name on 2026-07-06 — both predate Grok 4.6, and Grok 4.5 (July 2026) already shipped under this same combined corporate structure, so the "xAI (SpaceXAI)" lab label reflects the company's current name, not anything new about this release.
Not yet on tbench.ai, Toolathlon, or Snorkel boards — all predate its release; as of 2026-08-20 (8 days post-launch) none of the three had published a Grok 4.6 score.
Compare with
FAQ
Has Grok 4.6 been independently benchmarked?
Yes, on all nine metrics Grok 4.6 has scores for on this site. HLE, Terminal-Bench 2.1, SWE-bench Verified, GPQA Diamond, LiveCodeBench, DeepSWE, ARC-AGI-2, LiveBench, and AA-AnalystAgent are each independent runs — from Vals.ai, Artificial Analysis, the DeepSWE leaderboard (DataCurve), ARC Prize, and LiveBench's own board, observed 2026-08-17 through 2026-09-29 — not figures self-reported by xAI. That's 9 of 9 independently sourced; see the table above for the individual figures.
Why is there no tbench.ai, Toolathlon, or Snorkel score for Grok 4.6?
Grok 4.6 released 2026-08-12, and none of those three agentic-benchmark boards had published a run for it as of 2026-08-20. All three boards predate the model's release, so the gap reflects timing, not a refused or failed evaluation.
Does Grok 4.6's price change with longer prompts?
Yes. The $2 input / $6 output per-million-token rate shown on this page applies only to prompts under 200,000 tokens. At 200,000 tokens and above, xAI bills the entire request at the higher tier — $4 input / $12 output per million tokens, per docs.x.ai/developers/models (verified 2026-08-20). Cached input pricing doubles the same way, from $0.50 to $1.00 per million tokens, at the same 200k cutoff.
What reasoning effort levels does Grok 4.6 support?
Four settings — low, medium, high (the default), and xhigh — confirmed on xAI's reasoning docs (docs.x.ai/developers/model-capabilities/text/reasoning, verified 2026-08-20); xhigh is only available on Grok 4.6 and later, so requests for it on Grok 4.5 fall back to high. Higher settings trade latency and token spend for more deliberate reasoning; the benchmark scores on this page reflect whichever effort tier the source evaluation used, not necessarily the default.
Why does this page list Grok 4.6's lab as "xAI (SpaceXAI)"?
SpaceX acquired xAI in an all-stock merger on 2026-02-02, and xAI's own channels rebranded to "SpaceXAI" on 2026-07-06 (source: x.ai/news/xai-joins-spacex, widely reported by Reuters/Bloomberg/CNBC). Both changes predate Grok 4.6: Grok 4.5 (July 2026) was already released under the same combined corporate structure, so Grok 4.6 is not the first model shipped this way — the "xAI (SpaceXAI)" label here just reflects the lab's current corporate name, not anything new about this specific release.
What is Grok 4.6's knowledge cutoff?
February 2026, at month precision, with a 500,000-token context window. Two specs worth pairing with it: prompts at or above 200,000 tokens switch to a higher price tier, and the window is the number to check before any long-context work on this model.
Further reading
- GPQA Diamond leaderboard 2026 — Grok 4.6 is one of the 23 models it compares.
- Models with 10M token context windows 2026 — Grok 4.6 is one of the 35 models it compares.
Benchmark guides
- Humanity's Last Exam — what Grok 4.6’s reasoning score on it does and doesn’t prove.