Google DeepMind
Gemini 4 Argon
Announced September 30, not yet shipped: no API identifier, no release date, and the $2/$10 introductory rate doubles to $4/$20. Of this site's twelve tracked benchmarks, one has an independent score for it: 57.1 on Humanity's Last Exam at Artificial Analysis's high tier — 9.3 points above Gemini 3.8 Flash at that same tier, and the only row on this board not recorded at max effort.
Gemini 4 Argon benchmarks and pricing, every number sourced: 1 tracked Gemini 4 Argon benchmark score (1 independently run, 0 still resting on a vendor’s own claim), priced at $4.00 per million input tokens and $20.00 per million output.
Gemini 4 Argon’s 1 benchmark score on this page was verified against its source on 2026-10-02.
- Released
- 2026-09-30
- Availability
- Announced, not released: no API identifier and no general-release date as of 2026-10-02
- License
- proprietary
- Context window
- 1M tokensnot vendor-stated — Artificial Analysis's figure; Google publishes a 1M output limit, not a context window
- Knowledge cutoff
- Not disclosed
- Parameters
- Not disclosed
- Architecture
- Not disclosed
Gemini 4 Argon’s verified record
Gemini 4 Argon’s most-compared rival is Claude Fable 5.1: 1 trail across their 1 shared comparison. Gemini 4 Argon is announced at $4.00/$20.00 per 1M tokens (in/out) vs Claude Fable 5.1’s $10.00/$50.00.
Against the 32 head-to-head comparisons Gemini 4 Argon shares with other tracked models: 29 real gaps, 1 inside the noise band, and 2 we will not call.
A gap counts for Gemini 4 Argon only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Gemini 4 Argon trails on 2 of them.
HLE · no tools Reasoning
±2 is noiseGemini 4 Argon announced pricing
$4.00 in / $20.00 out per 1M tokens — Google's announcement
What Gemini 4 Argon costs per job
| Workload | Tokens in / out | Cost |
|---|---|---|
| One long chat turn | 100K / 10K | $0.600 |
| A codebase review | 1,000K / 100K | $6.00 |
| A day of agent work | 10,000K / 1,000K | $60.00 |
Computed from Gemini 4 Argon’s list rates above — cache discounts, its introductory rate, and batch tiers are not applied.
Gemini 4 Argon is one of 4 Google DeepMind models tracked on this site, at these official list prices.
| Model | In / 1M | Out / 1M |
|---|---|---|
| Gemini 4 Argon | $4.00 | $20.00 |
| Gemini 3.8 Flash | $0.75 | $3.75 |
| Gemini 3.7 Flash(superseded) | $0.75 | $3.75 |
| Gemini 3.1 Pro Preview | $2.00 | $12.00 |
Gemini 4 Argon benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise |
Who ran these numbers: 1 of 1 independent — artificialanalysis.ai (1).
- HLE: At AA's high effort tier, the only tier AA publishes for it. Read from the model page's own embedded record: Terminal-Bench 2.1, GPQA, LiveCodeBench and AA-AnalystAgent are all null there; Terminal-Bench 4.0 reads 57.07, which this site does not track.
Notes on the record
Google announced Gemini 4 Argon on 2026-09-30 (post published 20:00 UTC) and it is not generally available: no API model identifier appears in Google's own model list, pricing page, Vertex AI list, changelog or release notes, all checked 2026-10-01, and no general-release date is given. Access starts with trusted cyber defenders through the Fairwind Program; Google says it is engaged in the U.S. voluntary process for pre-release model access, and that defenders will get the model without cyber guardrails. Price is introductory $2/$10 per 1M tokens, after which it is $4/$20; Google states cached input as 95% off the input rate and never prints the resulting figure, so the $0.10 Artificial Analysis carries is arithmetic. No batch rate, tier or long-context surcharge is published. No knowledge cutoff and no reasoning-effort control is documented.
Google states a 1M output token limit, "up from the previous 64K tokens," which is an output ceiling and not a context window. The 1,000,000 recorded here is Artificial Analysis's context figure; no vendor-stated window exists. The output ceiling implies the window is at least 1M, since a model cannot emit more than it accepts, but Google has not said what the window is.
Artificial Analysis carries HLE 57.09% without tools, Intelligence Index 53 and Terminal-Bench 4.0 at 57.07, read from the model page's own embedded record where Argon appears once at effort high. vals.ai ranks it fifth on Terminal-Bench 4.0 at 57.58% (±2.31, Mini-SWE-agent) and first on its Vals Index at 68.90%. Terminal-Bench 4.0 is not a benchmark this site tracks. Google's own numbers are coding results it is not releasing publicly, including DeepSWE v1.1 at 77.9% as new state of the art; the official DeepSWE board does not list the model. No supersession language appears in the post and Gemini 3.1 Pro is never mentioned, so nothing tracked here is marked superseded.
Compare with
FAQ
Can you buy Gemini 4 Argon yet?
Not through a normal API. Google published an announcement on September 30, 2026 but has published no API model identifier, and the model list on its own developer site still ends at Gemini 3.8 Flash. Access starts with trusted cyber defenders through the Fairwind Program, with general release promised to paid API customers and Google AI Ultra subscribers "as soon as possible" and no date attached.
What does Gemini 4 Argon actually score on independent benchmarks?
One tracked figure: Artificial Analysis puts it at 57.09% on Humanity's Last Exam without tools, at its high effort tier. Two labs have run it on benchmarks this site does not track — vals.ai's own Terminal-Bench 4.0 run lists it fifth at 57.58% (±2.31) and first on its Vals Index at 68.90% — while LiveBench, DeepSWE, ARC-AGI, Toolathlon, Agents' Last Exam and HMMT were each checked and none lists it. Google's own coding claims are not on any of those boards.
Is the $2 per million token price going to stay?
No — that is the introductory rate. Google's post states $2 per million input tokens and $10 per million output to start, and $4 and $20 once the introductory period expires. Google gives no end date for that period. Cached input is priced at 95% off the input rate — $0.10 today, and $0.20 once the standard rate applies. Google publishes only the percentage, so both figures are arithmetic.
How does Argon compare with Google's earlier flagship Gemini models?
Artificial Analysis scores Gemini 3.8 Flash at 47.8% on the same Humanity's Last Exam variant and at the same high effort tier, so Argon is 9.3 points ahead against a 2-point noise band — a real gap on that evidence. Argon is also one of several rows here recorded below max effort, the 3.7 and 3.8 Flash siblings among them, so its place in that column is not a matched ranking. Google makes no supersession claim: its post never mentions Gemini 3.1 Pro, and names only 3.8 Flash Cyber as a predecessor.
What is Gemini 4 Argon good at?
The evidence is thin and mostly Google's own. It reports first on AutomationBench at 51.3%, tied for first on CWE-bench v1 at 68%, and state of the art on LVBench at 91.7%, plus 77.9% on DeepSWE v1.1 which it calls a new state of the art — the official DeepSWE board, 28 models and updated September 22, does not list it. The one independent measurement, HLE, is a reasoning-and-knowledge score rather than a coding one.
Further reading
- Models with 10M token context windows 2026 — Gemini 4 Argon is one of the 35 models it compares.