Google DeepMind
Gemini 3.1 Pro Preview
Google's own pricing page confirms the tiered $2/$12 (≤200K tokens) to $4/$18 (>200K) rates, and eleven of its twelve tracked benchmark scores are independently verified. The one figure DeepMind chose to headline — Humanity's Last Exam with tools, 51.4% — is still self-reported, and six months after launch the model remains in Preview, with DeepMind's own site already teasing a successor.
Gemini 3.1 Pro Preview benchmarks and pricing, every number sourced: 12 tracked Gemini 3.1 Pro Preview benchmark scores (11 independently run, 1 still resting on a vendor’s own claim), priced at $2.00 per million input tokens and $12.00 per million output.
Gemini 3.1 Pro Preview’s 12 benchmark scores on this page were each verified against their sources between 2026-07-01 and 2026-09-29.
- Released
- 2026-02-19
- License
- proprietary
- Context window
- 1M tokens
- Knowledge cutoff
- Not disclosed
- Parameters
- Not disclosed
- Architecture
- Not disclosed
Gemini 3.1 Pro Preview’s verified record
Gemini 3.1 Pro Preview’s featured comparison is Kimi K3: 2 trails, 3 ties, and 6 not callable. Gemini 3.1 Pro Preview is priced at $2.00/$12.00 per 1M tokens (in/out) vs Kimi K3’s $3.00/$15.00. Full verdict →
Against the 207 head-to-head comparisons Gemini 3.1 Pro Preview shares with other tracked models: 60 real gaps, 31 inside the noise band, and 116 we will not call.
A gap counts for Gemini 3.1 Pro Preview only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Gemini 3.1 Pro Preview trails on 40 of them.
HLE · no tools Reasoning
±2 is noiseLiveBench Composite score across 7 domains
±2.7 is noiseAgents' Last Exam Professional work
±3.2 is noiseAnalystAgent Spreadsheet & document analysis
±11.2 is noiseToolathlon-Verified Multi-tool chores
±9.7 is noiseTerminal-Bench 2.1 Terminal ops
±10.6 is noiseARC-AGI-2 · untiered Compositional visual reasoning
±9.2 is noiseNo verdict for Gemini 3.1 Pro Preview anywhere on GPQA Diamond, LiveCodeBench, SWE-bench Verified (saturated); HLE · with tools (nothing independently confirmed on both sides); HMMT (contaminated).
- None of Gemini 3.1 Pro Preview’s coding comparisons are independently confirmed on both sides yet.
Gemini 3.1 Pro Preview API pricing
$2.00 in / $12.00 out per 1M tokens — official pricing
What Gemini 3.1 Pro Preview costs per job
| Workload | Tokens in / out | Cost |
|---|---|---|
| One long chat turn | 100K / 10K | $0.320 |
| A codebase review | 1,000K / 100K | $3.20 |
| A day of agent work | 10,000K / 1,000K | $32.00 |
Computed from Gemini 3.1 Pro Preview’s list rates above — cache discounts and batch tiers are not applied.
Gemini 3.1 Pro Preview is one of 4 Google DeepMind models tracked on this site, at these official list prices.
| Model | In / 1M | Out / 1M |
|---|---|---|
| Gemini 4 Argon | $4.00 | $20.00 |
| Gemini 3.8 Flash | $0.75 | $3.75 |
| Gemini 3.7 Flash(superseded) | $0.75 | $3.75 |
| Gemini 3.1 Pro Preview | $2.00 | $12.00 |
Gemini 3.1 Pro Preview benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise | |
| HLE(with tools)[2] Reasoning · ±2 is noise | |
| GPQA Diamondsaturated[3] Expert science Q&A — not ranked at any gap size | |
| SWE-bench Verifiedsaturated[4] Bug fixing — not ranked at any gap size | |
| Terminal-Bench 2.1[5] Terminal ops · ±10.6 is noise | |
| Toolathlon-Verified Multi-tool chores · ±9.7 is noise | |
| Agents' Last Exam[6] Professional work · ±3.2 is noise | |
| LiveCodeBenchsaturated[7] Contest coding — not ranked at any gap size | |
| ARC-AGI-2(untiered)[8] Compositional visual reasoning · ±9.2 is noise | |
| LiveBench[9] Composite score across 7 domains · ±2.7 is noise | |
| HMMTcontaminated[10] Competition mathematics — not ranked at any gap size | |
| AnalystAgent[11] Spreadsheet & document analysis · ±11.2 is noise |
Who ran these numbers: 11 of 12 independent — vals.ai (3), artificialanalysis.ai (2), tbench.ai (1), toolathlon.xyz (1), snorkel.ai (1), arcprize.org (1), livebench.ai (1), matharena.ai (1); vendor self-reported (1).
- HLE: AA's own run (text-only subset, no tools). Replaces Google's self-reported 44.4 — independent beats vendor at the same setup.
- HLE: Reported as 'Search + Code tools', a superset of plain tool access.
- GPQA Diamond: vals.ai run, rank 1/133 (updated 2026-08-15). Replaces Google's self-reported 94.3. vals.ai notes 24 models at 90%+ — benchmark near saturation.
- SWE-bench Verified: vals.ai run, rank 24/83, bash-only mini-swe-agent harness (updated 2026-08-14). Replaces Google's self-reported 80.6.
- Terminal-Bench 2.1: Re-verified 2026-09-03 in a browser (found while investigating the check-sources MISS on this benchmark's Claude rows — this row cites the same now-stale deep link, but happened to pass the automated check on a coincidental substring match rather than a real one; see the claude-opus-4-8 row on this benchmark for the site-redesign mechanism). Value confirmed unchanged: 65.8%, rank 14 of 17, 'Gemini 3.1 Pro (high)' via Gemini CLI (CI now shown as ±3.3%). A Terminus-2-harness run on the same leaderboard scores 65.6 at rank 16 — essentially the same. As of 2026-08-17 this was the only one of this site's 5 then newly-added models with an independent Terminal-Bench 2.1 submission; the other four had only vendor-reported numbers 15-23 points higher. All four now carry independent rows. Cross-check (2026-10-01): vals.ai's archived Terminus 2 table lists 'Gemini 3.1 Pro Preview (02/26)' at 70.79.
- Agents' Last Exam: Overall pass rate 16.4 (Gemini CLI, High; score 32.7; $2,018). Corrects an earlier version of this page, which showed 32.7 — that was the partial-credit score, not the pass rate.
- LiveCodeBench: vals.ai run, rank 4/138 (updated 2026-08-15).
- ARC-AGI-2: Gemini 3.1 Pro Preview's official ARC-AGI-2 leaderboard row, dated 2026-02-19 on arcprize.org. This row carries no tier suffix — a flat, untiered score; a same-named 2026-03-05 row is a distinct entry with ARC-AGI-2 = N/A and is not a competing citation.
- LiveBench: Board row "Gemini 3.1 Pro Preview High" on the 2026-06-25 LiveBench release.
- HMMT: MathArena's leaderboard row is "Gemini 3.1 Pro Preview", 94.70% across all 33 HMMT Feb 2026 problems (live-verified 2026-08-25 at matharena.ai/?comp=hmmt--hmmt_feb_2026). Carries MathArena's own contamination-warning flag — "Model was released after competition release" — a disclosed risk that applies to every one of this site's 4 tracked HMMT Feb 2026 rows, Gemini 3.1 Pro Preview included, since all 4 tracked models were released after HMMT Feb 2026 took place.
- AnalystAgent: AA's own run, board row 'Gemini 3.1 Pro Preview' (untiered on AA); the board rounds to one decimal. Gemini 3.1 Pro Preview's value was read from AA's model page on 2026-09-29; the leaderboard's embedded payload carries only its 14 default-selected models, which excludes this one.
Notes on the record
Pricing shown is for prompts <=200K context; rises to $4/$18 per 1M tokens beyond that (confirmed on Google's official pricing page as of 2026-08-20). A Batch API tier is also available at roughly half price: $1 input / $6 output per 1M tokens for prompts <=200K, and $2/$9 above 200K (source: ai.google.dev/gemini-api/docs/pricing).
Still listed as Preview — Google DeepMind's own model page (deepmind.google/models/gemini/pro/) still showed "Preview" status as of 2026-08-20, more than six months after the 19 February 2026 launch, and that same page already references a "3.5 Pro" successor as coming soon. Output is capped at 65,536 tokens per response, separate from the 1,048,576-token input context window (source: DeepMind model card).
Published 19 February 2026 per the official DeepMind model card; that card states no knowledge cutoff, so this field stays empty. Note: Google's own Gemini API developer documentation (ai.google.dev/gemini-api/docs/gemini-3) lists a models table entry for gemini-3.1-pro-preview showing a knowledge cutoff of "Jan 2025" -- the same domain this page already cites as its pricing source -- and Google Cloud's Vertex AI documentation shows the same January 2025 date. Neither detail appears on the primary DeepMind model card this site sources from for the knowledge_cutoff field -- flagged for human review rather than silently added to the data.
Compare with
FAQ
Has Gemini 3.1 Pro Preview been independently benchmarked?
Mostly, yes. Eleven of the twelve benchmark scores tracked on this page for Gemini 3.1 Pro Preview come from independent evaluators — Artificial Analysis (Humanity's Last Exam, no tools, and AA-AnalystAgent), Vals.ai (GPQA Diamond, SWE-bench Verified, LiveCodeBench), Terminal-Bench, Toolathlon, Snorkel AI's Agents' Last Exam, ARC Prize (ARC-AGI-2), LiveBench's own board, and MathArena (HMMT Feb 2026) — observed between 2026-07-01 and 2026-09-29 per the table above. The exception is the Humanity's Last Exam with-tools score (51.4%, per Google DeepMind's 19 February 2026 model card), which is self-reported and has not been reproduced by an independent board as of 2026-08-20.
Does Gemini 3.1 Pro Preview's API price change with context length?
Yes. Google's Gemini API pricing page lists $2 per 1M input tokens and $12 per 1M output tokens for Gemini 3.1 Pro Preview at prompts up to 200K tokens, rising to $4 input / $18 output per 1M tokens once a prompt exceeds 200K tokens. Since the model's context window runs to 1,048,576 tokens, a long-document or full-repo prompt to Gemini 3.1 Pro Preview can land well into the higher tier. Source: ai.google.dev/gemini-api/docs/pricing, checked 2026-08-20.
Is Gemini 3.1 Pro Preview still labeled Preview, or has it reached general availability?
Still Preview, per this page's 2026-08-20 check. Google DeepMind's own model listing for Gemini 3.1 Pro shows a 'Preview' status more than six months after its 19 February 2026 launch, and the same page already references a '3.5 Pro' model as coming soon — so the Preview label for Gemini 3.1 Pro Preview may persist right up until it's superseded rather than resolving into a GA release.
Is there a lower-cost way to run Gemini 3.1 Pro Preview?
Yes — Google's Batch API prices Gemini 3.1 Pro Preview at roughly half the standard synchronous rate: $1 input / $6 output per 1M tokens for prompts up to 200K, and $2 input / $9 output above 200K, versus $2/$12 and $4/$18 on the standard API. The tradeoff is that batch requests are processed asynchronously rather than returned immediately, so it suits bulk/offline workloads rather than interactive use. Source: ai.google.dev/gemini-api/docs/pricing, checked 2026-08-20.
Why doesn't this page list a knowledge cutoff for Gemini 3.1 Pro Preview?
Google DeepMind's official model card for Gemini 3.1 Pro Preview (published 19 February 2026) does not state a knowledge cutoff date, so this site leaves the field empty rather than guess. For reference, Google's own Gemini API developer documentation (ai.google.dev/gemini-api/docs/gemini-3) lists a 'Jan 2025' knowledge cutoff for gemini-3.1-pro-preview in its models table, and Vertex AI's documentation shows the same date — a detail absent from the primary DeepMind model card this site sources from, and not yet reconciled here.
Further reading
- GPQA Diamond leaderboard 2026 — Gemini 3.1 Pro Preview is one of the 23 models it compares.
- Models with 10M token context windows 2026 — Gemini 3.1 Pro Preview is one of the 35 models it compares.