xAI (SpaceXAI)
Grok 4.7
xAI's newest flagship, priced and speed-matched to Grok 4.6. Independent coverage is thin so far — just Humanity's Last Exam (43.1% at xhigh effort, versus Grok 4.6's 42.9% at high effort, not a same-tier comparison) and LiveBench (77.4, essentially level with Grok 4.6's 78.0). Its DeepSWE figure is a vendor claim, not a board run, and no GPQA Diamond, SWE-bench, LiveCodeBench, ARC-AGI-2 or Terminal-Bench 2.1 score exists yet.
Grok 4.7 benchmarks and pricing, every number sourced: 3 tracked Grok 4.7 benchmark scores (2 independently run, 1 still resting on a vendor’s own claim), priced at $2.00 per million input tokens and $6.00 per million output.
Grok 4.7’s 3 benchmark scores on this page were verified against their source on 2026-09-23.
- Released
- 2026-09-21
- License
- proprietary
- Context window
- 500K tokens
- Knowledge cutoff
- 2026-05
- Parameters
- Not disclosed
- Architecture
- Not disclosed
Grok 4.7’s verified record
Grok 4.7’s most-compared rival is GLM-5.3-Flash: 2 leads and 1 not callable across their 3 shared comparisons. Grok 4.7 is priced at $2.00/$6.00 per 1M tokens (in/out) vs GLM-5.3-Flash’s $0.15/$0.50.
Against the 77 head-to-head comparisons Grok 4.7 shares with other tracked models: 34 real gaps, 17 inside the noise band, and 26 we will not call.
A gap counts for Grok 4.7 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Grok 4.7 trails on 21 of them.
HLE · no tools Reasoning
±2 is noiseLiveBench Composite score across 7 domains
±2.7 is noiseNo verdict for Grok 4.7 anywhere on DeepSWE (nothing independently confirmed on both sides).
- None of Grok 4.7’s coding comparisons are independently confirmed on both sides yet.
Grok 4.7 API pricing
$2.00 in / $6.00 out per 1M tokens — official pricing
What Grok 4.7 costs per job
| Workload | Tokens in / out | Cost |
|---|---|---|
| One long chat turn | 100K / 10K | $0.260 |
| A codebase review | 1,000K / 100K | $2.60 |
| A day of agent work | 10,000K / 1,000K | $26.00 |
Computed from Grok 4.7’s list rates above — cache discounts and batch tiers are not applied.
Grok 4.7 is one of 2 xAI (SpaceXAI) models tracked on this site, at these official list prices.
| Model | In / 1M | Out / 1M |
|---|---|---|
| Grok 4.7 | $2.00 | $6.00 |
| Grok 4.6 | $2.00 | $6.00 |
Grok 4.7 benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise | |
| LiveBench[2] Composite score across 7 domains · ±2.7 is noise | |
| DeepSWE[3] Long-horizon coding · ±9.5 is noise |
Who ran these numbers: 2 of 3 independent — artificialanalysis.ai (1), livebench.ai (1); vendor self-reported (1).
- HLE: AA's own run at xhigh effort (Grok 4.6's tracked score used high effort) — text-only 2,158-question subset, part of Intelligence Index v4.3.2.
- LiveBench: Board row 'Grok 4.7 xHigh' on the same 2026-06-25 LiveBench release Grok 4.6 sits on; only the xHigh tier is populated. Grok 4.6 scores slightly higher overall on this table (78.0 vs 77.4).
- DeepSWE: xAI's own model card figure (DeepSWE v1.1, high effort). The independent deepswe.datacurve.ai board (checked live 2026-09-23, 28 tracked models) does not list this model yet — only grok-4.6 and grok-4.5 appear.
Notes on the record
Pricing follows the identical two-tier scheme as Grok 4.6, confirmed on docs.x.ai/docs/models (checked 2026-09-23): $2.00 input / $6.00 output per million tokens under 200,000 prompt tokens, rising to $4.00/$12.00 — the full request billed at the higher rate — once a prompt reaches 200,000. Cached input follows the same split, $0.50 rising to $1.00. A "Fast" variant bills at twice the standard rates; xAI's developer docs say it is available only in Cursor and Grok Build, not the public API.
xAI's developer docs list no text output limit. Knowledge cutoff is May 2026 per the docs, while the model card gives a June 2026 pretraining cutoff — both official, recorded unresolved. Parameter count is undisclosed in xAI's model card; widely circulated third-party reports of 2.1 trillion parameters trace to no xAI source found and are not used here. Input: text and images. Output: text only.
Independent coverage is thin so far. Grok 4.6 eventually reached independent scores across every benchmark this page tracks for it, but not at launch — those scores were observed 5 to 12 days after its 2026-08-12 release. Grok 4.7's record, two days in, is Humanity's Last Exam (Artificial Analysis, text-only, xhigh effort — a different tier from Grok 4.6's tracked high-effort score) and LiveBench; DeepSWE, GPQA Diamond, SWE-bench Verified, LiveCodeBench, ARC-AGI-2 and Terminal-Bench 2.1 have no independent score yet — deepswe.datacurve.ai and arcprize.org were checked directly and neither lists this model as of 2026-09-23. xAI does publish a self-reported DeepSWE v1.1 score (below) and a Terminal-Bench 4.0 score of 38.0%, a different, newer benchmark version from the Terminal-Bench 2.1 this site tracks, so no TB 2.1 figure exists here.
xAI's own materials never say Grok 4.7 supersedes or replaces Grok 4.6 — Grok 4.6 stays listed on the same pricing table, not flagged deprecated. "xAI (SpaceXAI)" reflects the same July 2026 corporate rebrand already explained on the Grok 4.6 page.
Compare with
FAQ
Has Grok 4.7 been independently benchmarked?
Only partly, as of 2026-09-23. Artificial Analysis has an independent Humanity's Last Exam score (43.1%, text-only, xhigh effort) and LiveBench has an independent board row (77.4 overall, xHigh). Neither GPQA Diamond, SWE-bench Verified, LiveCodeBench, ARC-AGI-2, Terminal-Bench 2.1 nor the independent DeepSWE leaderboard lists this model yet — deepswe.datacurve.ai and arcprize.org were checked directly and found no Grok 4.7 row. Grok 4.6 eventually reached independent scores across every benchmark this page tracks for it too, but that took 5 to 12 days after its own launch, not day one.
Does Grok 4.7's price change with longer prompts?
Yes, the same way Grok 4.6's does. The $2 input / $6 output per-million-token rate applies under 200,000 prompt tokens; at 200,000 and above, xAI bills the whole request at $4 input / $12 output, per docs.x.ai/docs/models (checked 2026-09-23). Cached input doubles the same way, $0.50 to $1.00.
How many parameters does Grok 4.7 have?
xAI has not disclosed a parameter count in its official model card or documentation. Third-party reports circulating online cite 2.1 trillion parameters, but no xAI source states this figure, so it isn't used on this page.
Does Grok 4.7 replace Grok 4.6?
Not according to xAI's own materials. Grok 4.6 remains listed on the same pricing page, with no deprecation notice, and xAI's announcement and model card describe Grok 4.7 as extending Grok 4.6's capability rather than retiring it — "served at the same price and speed as Grok 4.6," per xAI's own launch post.
Is Grok 4.7's Terminal-Bench score comparable to other models on this site?
No. xAI reports a self-reported Terminal-Bench 4.0 score of 38.0%, a different, newer benchmark version from the Terminal-Bench 2.1 this site tracks for other models. No Terminal-Bench 2.1 score exists for Grok 4.7.
Further reading
- Models with 10M token context windows 2026 — Grok 4.7 is one of the 31 models it compares.