Google DeepMindPrevious version · Gemini 3.8 Flash

Gemini 3.7 Flash

The $0.75/$3.75 launch price is real, and all nine benchmark scores on this page are independent, third-party runs — it's Google's own listed rate and every score is independently verified — but that price is a four-and-a-half-month introductory offer that doubles to $1.50/$7.50 the moment 2027 starts.

Gemini 3.7 Flash benchmarks and pricing, every number sourced: 9 tracked Gemini 3.7 Flash benchmark scores (9 independently run, 0 still resting on a vendor’s own claim), priced at $0.75 per million input tokens and $3.75 per million output.

Released
2026-08-13
License
proprietary
Context window
1M tokens
Knowledge cutoff
2026-03
Verified
sources checked 2026-08-17–2026-09-29
Parameters
Not disclosed
Architecture
Not disclosed

Gemini 3.7 Flash’s verified record

Gemini 3.7 Flash’s most-compared rival is Qwen3.8-Max: 2 leads, 2 ties, and 4 not callable across their 8 shared comparisons. Gemini 3.7 Flash is priced at $0.75/$3.75 per 1M tokens (in/out) vs Qwen3.8-Max’s $2.00/$6.00.

Against the 199 head-to-head comparisons Gemini 3.7 Flash shares with other tracked models: 23 real gaps, 25 inside the noise band, and 151 we will not call.

A gap counts for Gemini 3.7 Flash only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Gemini 3.7 Flash trails on 1 of them.

HLE · no tools Reasoning

±2 is noise
Behind
Ahead
Qwen3.8-Max +4.8 · DeepSeek V4 Pro (0813) +6.9 · GLM-5.3-Flash +8.0 · MiniMax M3 +8.9 · Qwen3.8-Flash-Next +9.9 · 2 previous: Grok 4.6 +5.0 · GLM-5.2 +6.8
Tie
6 models within ±2
Unverified
1 model — vendor-reported on one side
Setup-dependent
18 models — scored on a different harness or effort tier

LiveBench Composite score across 7 domains

±2.7 is noise
Ahead
GLM-5.3 +2.7 · Gemini 3.8 Flash +3.0 · GLM-5.3-Flash +7.2 · MiniMax M3 +11.5 · 2 previous: DeepSeek V4 Flash (0731) +4.6 · GLM-5.2 +5.6
Tie
6 models within ±2.7
Setup-dependent
16 models — scored on a different harness or effort tier

AnalystAgent Spreadsheet & document analysis

±11.2 is noise
Ahead
Qwen3.8-Max +15.0 · Gemini 3.1 Pro Preview +18.8 · Step 5 Preview +25.0 · MiniMax M3 +50.0 · 1 previous: Grok 4.6 +18.8
Setup-dependent
11 models — scored on a different harness or effort tier

DeepSWE Long-horizon coding

±9.5 is noise
Ahead
3 previous: Claude Sonnet 5 +11.0 · DeepSeek V4 Flash (0731) +12.0 · GLM-5.2 +21.0
Tie
7 models within ±9.5
Unverified
11 models — vendor-reported on one side
Setup-dependent
8 models — scored on a different harness or effort tier

ARC-AGI-2 · high Compositional visual reasoning

±9.2 is noise
Ahead
1 previous: Claude Opus 4.8 +12.5
Tie
2 models within ±9.2

No verdict for Gemini 3.7 Flash anywhere on Terminal-Bench 2.1 ( every independently confirmed comparison inside the noise band); GPQA Diamond, LiveCodeBench, SWE-bench Verified ( saturated).

Gemini 3.7 Flash API pricing

$0.75 in / $3.75 out per 1M tokens — official pricing source

What Gemini 3.7 Flash costs per job

Gemini 3.7 Flash cost for three reference workloads, computed from its list rates
WorkloadTokens in / outCost
One long chat turn100K / 10K$0.113
A codebase review1,000K / 100K$1.13
A day of agent work10,000K / 1,000K$11.25

Computed from Gemini 3.7 Flash’s list rates above — cache discounts and batch tiers are not applied.

Gemini 3.7 Flash is one of 4 Google DeepMind models tracked on this site, at these official list prices.

Google DeepMind model pricing, official list rates
ModelIn / 1MOut / 1M
Gemini 4 Argon$4.00$20.00
Gemini 3.8 Flash$0.75$3.75
Gemini 3.7 Flash(previous version)$0.75$3.75
Gemini 3.1 Pro Preview$2.00$12.00

Gemini 3.7 Flash benchmark scores

Gemini 3.7 Flash benchmark scores, provenance, and source links
BenchmarkScore
HLE(no tools)[1]
Reasoning · ±2 is noise
SWE-bench Verifiedsaturated[2]
Bug fixing — not ranked at any gap size
GPQA Diamondsaturated[3]
Expert science Q&A — not ranked at any gap size
LiveCodeBenchsaturated[4]
Contest coding — not ranked at any gap size
DeepSWE[5]
Long-horizon coding · ±9.5 is noise
Terminal-Bench 2.1[6]
Terminal ops · ±10.6 is noise
ARC-AGI-2(high)[7]
Compositional visual reasoning · ±9.2 is noise
LiveBench[8]
Composite score across 7 domains · ±2.7 is noise
AnalystAgent[9]
Spreadsheet & document analysis · ±11.2 is noise

Who ran these numbers: 9 of 9 independent — artificialanalysis.ai (3), vals.ai (3), deepswe.datacurve.ai (1), arcprize.org (1), livebench.ai (1).

  1. HLE: AA's own run of Gemini 3.7 Flash at high effort, text-only subset. Beats Grok 4.6 here despite a lower overall AA index — single benchmarks and composites disagree, which is rather the point of this site.
  2. SWE-bench Verified: vals.ai run, bash-only harness (updated 2026-08-14).
  3. GPQA Diamond: vals.ai run (93.94), rank 4 (updated 2026-08-15). Corrected 2026-10-03: the value field read 93.9 (truncated); the board displays 93.94.
  4. LiveCodeBench: vals.ai run (88.652), rank 3 — just ahead of Grok 4.6 (updated 2026-08-15). Corrected 2026-10-03: the value field read 88.7 (mis-rounded); the board displays 88.65.
  5. DeepSWE: 65%±2 on mini-swe-agent, high effort (board updated 2026-08-13).
  6. Terminal-Bench 2.1: AA's own run of 'Gemini 3.7 Flash (high)'. Not on tbench.ai's official board, which lists only Gemini 3 Pro and 3.1 Pro. Cross-check (2026-10-01): vals.ai's archived Terminus 2 table lists it at 77.53.
  7. ARC-AGI-2: Gemini 3.7 Flash's official ARC-AGI-2 leaderboard row, dated 2026-08-13 on arcprize.org. No tier above High exists for this model — the top of three populated tiers (High/Medium/Low).
  8. LiveBench: Board row "Gemini 3.7 Flash High" on the 2026-06-25 LiveBench release.
  9. AnalystAgent: AA's own run, board row 'Gemini 3.7 Flash (high)' — same high tier as this site's other AA rows for this model. pass@1 70.5, pass@5 77.5.

Notes on the record

Introductory pricing through 2026-12-31; standard rate from 2027-01-01 is $1.50/$7.50, per Google's Gemini API pricing docs (https://ai.google.dev/gemini-api/docs/pricing and https://ai.google.dev/gemini-api/docs/latest-model). Output price includes thinking tokens; Google's docs describe this as three selectable reasoning tiers (low/medium/high) rather than a raw thinking-token budget.

GA per the official Gemini API changelog (2026-08-13), superseding Gemini 3.6 Flash. Knowledge cutoff March 2026 per the official DeepMind model card, which warns coverage in some domains only reaches January 2025. Google's Pro line (3.1) is still in preview — the Flash line is its only GA track. The discounted introductory rate also applies to the older Gemini 3.6 Flash, per Google's pricing page.

Per DeepMind's model card, 3.7 Flash is built on the Gemini 3.6 Flash base with algorithmic reasoning improvements, not a new pretraining run. The GA release drops support for deprecated sampling parameters (temperature, top_p, top_k, and candidate_count) that worked on earlier Gemini versions, per Google's migration/changelog docs — existing integrations calling those fields need updating.

Compare with

FAQ

Has Gemini 3.7 Flash been independently benchmarked?

Yes — every one of Gemini 3.7 Flash's nine tracked scores is independently sourced. HLE, Terminal-Bench 2.1, SWE-bench Verified, GPQA Diamond, LiveCodeBench, DeepSWE, ARC-AGI-2, LiveBench, and AA-AnalystAgent are each independent third-party runs — via Artificial Analysis, Vals.ai, the DeepSWE leaderboard, ARC Prize, and LiveBench's own board — observed 2026-08-17 through 2026-09-29, with none of the headline numbers on this page sourced from Google's own launch materials.

How much does Gemini 3.7 Flash cost, and will the price change?

$0.75 per million input tokens and $3.75 per million output tokens (thinking tokens are billed as output) — but that is introductory pricing, good only through 2026-12-31. From 2027-01-01 the rate rises to $1.50 in / $7.50 out per million tokens, per Google's Gemini API pricing page. The same discounted rate currently also applies to the older Gemini 3.6 Flash.

Does Gemini 3.7 Flash have a generally available Pro-tier sibling?

No. As of August 2026, Google's Pro line (Gemini 3.1 Pro) is still shipping only as a preview offering, not GA. Per Google's own materials, the Flash line — the track Gemini 3.7 Flash belongs to — is currently Google's only generally-available Gemini line.

What is Gemini 3.7 Flash's knowledge cutoff?

Google's official DeepMind model card lists March 2026 as the nominal cutoff, but explicitly warns that in some domains actual coverage only reaches back to January 2025 — so treat 'March 2026' as an upper bound rather than a guarantee for any specific topic.

What breaks for developers upgrading from Gemini 3.6 Flash to Gemini 3.7 Flash?

Beyond the benchmark gains, Google's own migration documentation notes that Gemini 3.7 Flash drops support for older sampling parameters — temperature, top_p, top_k, and candidate_count — that worked on earlier Gemini versions, so any integration relying on those fields needs to be updated before switching over.

Which Gemini 3.7 Flash score should you actually read?

Terminal-Bench 2.1 at 85.8 or ARC-AGI-2 at 84.6 on a high-effort run — the two current-benchmark peaks in its verified record. The larger figures above them, GPQA Diamond 93.94 and LiveCodeBench 88.65, both live on boards this site no longer ranks.

Further reading

Benchmark guides