Google DeepMind

Gemini 3.8 Flash

Gemini 3.8 Flash launched at the identical $0.75/$3.75 introductory price as its own predecessor and, per Google's own model card, isn't a new pretraining run at all — it's Gemini 3.7 Flash with post-training improvements — yet six of the eleven benchmarks this site tracks already carry an independently-run score one day after launch, including a DeepSWE result (74% at high effort) that ties Claude Opus 5 for the top spot on that board.

Gemini 3.8 Flash benchmarks and pricing, every number sourced: 6 tracked Gemini 3.8 Flash benchmark scores (6 independently run, 0 still resting on a vendor’s own claim), priced at $0.75 per million input tokens and $3.75 per million output.

Gemini 3.8 Flash’s 6 benchmark scores on this page were verified against their source on 2026-09-03.

Version history: succeeded Gemini 3.7 Flash (2026-08-13).

Released
2026-09-02
License
proprietary
Context window
1M tokens
Knowledge cutoff
2026-03
Parameters
Not disclosed
Architecture
Not disclosed

Gemini 3.8 Flash’s verified record

Against the 115 head-to-head comparisons Gemini 3.8 Flash shares with other tracked models: 27 real gaps, 38 inside the noise band, and 50 we will not call.

A gap counts for Gemini 3.8 Flash only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Gemini 3.8 Flash trails on 4 of them.

HLE · no tools Reasoning

±2 is noise
Ahead
Behind
Claude Fable 5.1 11.3 · Claude Fable 5 7.7 · Claude Opus 5 7.1 · GPT-6 Astra 6.9
Tie
6 models within ±2

DeepSWE Long-horizon coding

±9.5 is noise
Ahead
GLM-5.2 +30.0 · DeepSeek V4 Flash (0731) +21.0 · Claude Sonnet 5 +20.0 · Muse Spark 1.2 +19.0 · Qwen3.8-Max +17.0 · Claude Opus 4.8 +15.0 · GLM-5.3-Flash +11.0
Tie
9 models within ±9.5
Unverified
2 models — vendor-reported on one side

LiveCodeBench Contest coding

±3.1 is noise
Ahead
GLM-5.2 +20.0 · GLM-5.3 +9.0 · MiniMax M3 +7.3 · Claude Sonnet 5 +7.0 · GPT-5.6 Sol +6.9
Tie
11 models within ±3.1

No verdict for Gemini 3.8 Flash anywhere on Terminal-Bench 2.1 (every independently confirmed comparison inside the noise band); GPQA Diamond, SWE-bench Verified (saturated).

  • No real gap yet in any of Gemini 3.8 Flash’s agentic comparisons — the independently confirmed ones all sit inside the noise band.

What changed from Gemini 3.7 Flash to Gemini 3.8 Flash

The 6 benchmarks both models have been scored on, using the same variant each time. A raw Gemini 3.8 Flash gain is not a real gain until it clears that benchmark’s own noise band, so each row below carries the verdict and not just the arithmetic.

Gemini 3.8 Flash versus Gemini 3.7 Flash, per-benchmark change and whether it clears the noise band
BenchmarkGemini 3.7 FlashGemini 3.8 FlashChangeVerdict
HLE47.947.8-0.1Tie
Terminal-Bench 2.185.887.6+1.8Tie
DeepSWE6574+9.0Tie
SWE-bench Verified80.880-0.8Tainted
GPQA Diamond93.994.44+0.5Tainted
LiveCodeBench88.789.48+0.8Tie

Gemini 3.8 Flash API pricing

$0.75 in / $3.75 out per 1M tokens official pricing

What Gemini 3.8 Flash costs per job

Gemini 3.8 Flash cost for three reference workloads, computed from its list rates
WorkloadTokens in / outCost
One long chat turn100K / 10K$0.113
A codebase review1,000K / 100K$1.13
A day of agent work10,000K / 1,000K$11.25

Computed from Gemini 3.8 Flash’s list rates above — cache discounts and batch tiers are not applied.

Gemini 3.8 Flash is one of 3 Google DeepMind models tracked on this site, at these official list prices.

Google DeepMind model pricing, official list rates
ModelIn / 1MOut / 1M
Gemini 3.8 Flash$0.75$3.75
Gemini 3.7 Flash(superseded)$0.75$3.75
Gemini 3.1 Pro Preview$2.00$12.00

Gemini 3.8 Flash benchmark scores

Gemini 3.8 Flash benchmark scores, provenance, and source links
BenchmarkScore
HLE(no tools)[1]
Reasoning · ±2 is noise
Terminal-Bench 2.1[2]
Terminal ops · ±10.6 is noise
GPQA Diamondsaturated[3]
Expert science Q&A — not ranked at any gap size
LiveCodeBench[4]
Contest coding · ±3.1 is noise
SWE-bench Verifiedsaturated[5]
Bug fixing — not ranked at any gap size
DeepSWE[6]
Long-horizon coding · ±9.5 is noise

Who ran these numbers: 6 of 6 independent — vals.ai (3), artificialanalysis.ai (2), deepswe.datacurve.ai (1).

  1. HLE: AA's own run at high effort, text-only subset (raw payload value 0.478220574606117).
  2. Terminal-Bench 2.1: AA's own run of 'Gemini 3.8 Flash (high)' (0.876404494382023 in AA's payload); AA also lists medium (83.9%) and low (83.15%) effort tiers, not tracked here. Not yet on tbench.ai's official board. vals.ai's own harness (Terminus 2) scores it lower still, 81.27% ±0.38 — a 6.3-point spread between two independent sources on the same benchmark.
  3. GPQA Diamond: vals.ai run (94.44%), rank 4 of 138.
  4. LiveCodeBench: vals.ai run (89.48%), rank 3.
  5. SWE-bench Verified: vals.ai run, bash-only mini-swe-agent harness (80.00%±1.79), rank 24 of 88 — fractionally below Gemini 3.7 Flash's own 80.8% on the same board, a difference well inside this benchmark's own noise threshold.
  6. DeepSWE: 74%±1 on mini-swe-agent, high effort (v1.1 board, updated 2026-09-02) — same high-effort tier as Gemini 3.7 Flash's own 65% row on the same board (a 9-point gap that still lands just inside this benchmark's meaningful_gap of 9.5); ties Claude Opus 5 [max] for the top spot.

Notes on the record

Gemini 3.8 Flash launched September 2, 2026, per Google's own announcement blog post and DeepMind model card (independently corroborated by 9to5Google and Android Headlines the same day). Per its own model card, it "is based on Gemini 3.7 Flash" — an incremental update, not a new pretraining run — which is why its knowledge cutoff carries over unchanged (see below). Google's own comparison table treats it as the next iteration built on Gemini 3.7 Flash; that model's own entry on this site is marked superseded accordingly.

Headline API pricing is $0.75/MTok input and $3.75/MTok output — identical to Gemini 3.7 Flash's own current introductory rate, both confirmed side by side on Google's Gemini API pricing page (checked 2026-09-03). Both rates are introductory only: from 2027-01-01 they double to $1.50 in / $7.50 out. Batch API pricing is half the standard rate ($0.375/$1.875, doubling the same way on the same date), per the same page.

The context window carries over at 1,048,576 tokens in, per Google's Gemini API models reference; max output is 65,536 tokens (Google's own model card rounds this to "64K"). Knowledge cutoff is March 2026, with Google's own model card cautioning that some domains' coverage only reaches January 2025 — the identical wording and date used for Gemini 3.7 Flash's own cutoff, consistent with 3.8 Flash's own statement that it is not a new pretraining run.

A specialized sibling shipped the same day: Gemini 3.8 Flash Cyber, "powered by the same foundational intelligence" per Google's announcement but tuned for vulnerability detection and automated patching, restricted to vetted government and critical-infrastructure operators under Google's invitation-only Fairwind Program — the same access-restriction pattern this site has recorded for other labs' restricted-access siblings (e.g. Claude Mythos 5.1 — a different case in substance, the identical model with fewer guardrails rather than a tuned variant, but gated the same way). Cyber is not independently benchmarked on any of the eleven boards this site tracks and is not covered separately here.

One day after launch, six of the eleven benchmarks this site tracks already carry an independently-run Gemini 3.8 Flash score — GPQA Diamond, SWE-bench Verified, LiveCodeBench, Humanity's Last Exam, Terminal-Bench 2.1, and DeepSWE — all from vals.ai, Artificial Analysis, or DeepSWE's own board, none from Google's own claimed numbers. ARC-AGI-2 and LiveBench both scored Gemini 3.7 Flash but had not added Gemini 3.8 Flash as of 2026-09-03, checked live on each; Toolathlon-Verified, Agents' Last Exam, and HMMT Feb 2026 have never scored any Gemini Flash-line model, 3.8 Flash included.

Google's own launch materials claim an "HLE-Verified" score of 54.9% — a different figure from, and not directly comparable to, the 47.8% this site records from Artificial Analysis's independent no-tools run; Google discloses neither a methodology note nor a benchmark-family relationship for its own number, so the two are reported separately rather than merged.

DeepSWE's own v1.1 board scores Gemini 3.8 Flash at high effort (74%, tied with Claude Opus 5 for the top spot); Gemini 3.7 Flash's own tracked DeepSWE score (65%) used the same high-effort tier, so this is a clean same-effort comparison — the 9-point gap still lands just inside this benchmark's own noise floor (meaningful_gap 9.5 on 113 tasks), not a real gap. Terminal-Bench 2.1 is Artificial Analysis's own run (Terminus 2 harness, 87.6%); the benchmark's canonical board, tbench.ai, had not added Gemini 3.8 Flash as of this writing, and vals.ai's own harness (also called Terminus 2, but a different independent run) scores it lower, 81.27% ±0.38 — a 6.3-point spread between two sources this site treats as equally independent.

Not every number moved up: on SWE-bench Verified, Gemini 3.8 Flash's own vals.ai score (80.00%) sits fractionally below Gemini 3.7 Flash's own tracked score on the same board (80.80%) — a 0.8-point difference that is noise, not a regression, well inside this benchmark's own meaningful-gap threshold.

Compare with

FAQ

Has Gemini 3.8 Flash been independently benchmarked?

Partially so far — six of the eleven benchmarks this site tracks already carry an independently-run score, one day after the model's 2026-09-02 launch: GPQA Diamond (94.44%, vals.ai), SWE-bench Verified (80.00%, vals.ai), LiveCodeBench (89.48%, vals.ai), Humanity's Last Exam (47.8%, Artificial Analysis, no-tools), Terminal-Bench 2.1 (87.6%, Artificial Analysis), and DeepSWE (74% at high effort, tied for the board's top spot). ARC-AGI-2 and LiveBench both scored the prior Gemini 3.7 Flash but had not added 3.8 Flash as of 2026-09-03; Toolathlon-Verified, Agents' Last Exam, and HMMT Feb 2026 have never scored any Gemini Flash-line model. None of the six tracked scores here come from Google's own claimed numbers.

How does Gemini 3.8 Flash pricing compare to Gemini 3.7 Flash?

It's identical: $0.75 per million input tokens and $3.75 per million output tokens, confirmed side by side on Google's own Gemini API pricing page — the same introductory rate Gemini 3.7 Flash already carries, both good only through 2026-12-31. From 2027-01-01 both models' rates double to $1.50 in / $7.50 out per million tokens.

Is Gemini 3.8 Flash a new model, or an update to Gemini 3.7 Flash?

An update, by Google's own account: its model card states plainly that "Gemini 3.8 Flash is based on Gemini 3.7 Flash," pointing to the earlier model's own card for architecture and training-data details rather than describing a new pretraining run. That lineage is also why the knowledge cutoff carries over unchanged — March 2026, with the same caveat that some domains' coverage only reaches January 2025.

What is Gemini 3.8 Flash Cyber, and is it the same model?

A separate, restricted release announced the same day. Google describes it as "powered by the same foundational intelligence" as the base Gemini 3.8 Flash but tuned specifically for vulnerability detection and automated patching, and it is limited to vetted government agencies and critical-infrastructure operators under Google's invitation-only Fairwind Program — a similar access restriction to other labs' restricted-access siblings this site has recorded elsewhere, such as Claude Mythos 5.1 (which, unlike Cyber, is the identical model with fewer guardrails rather than a tuned variant). Cyber is not independently benchmarked on any board this site tracks and isn't covered separately here.

Does Gemini 3.8 Flash replace Gemini 3.7 Flash, or sit alongside it?

It replaces it. Google's own materials describe 3.8 Flash as the next iteration built on 3.7 Flash, the same relationship this site uses to mark a model superseded (as with Claude Fable 5 → Claude Fable 5.1 and GLM-5.2 → GLM-5.3) — Gemini 3.7 Flash stays on this site for comparison, not deleted, but is no longer the current recommendation.

Further reading

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.