DeepSeek

DeepSeek V4 Pro (0813)

DeepSeek V4 Pro (0813)'s coding and no-tools reasoning scores (SWE-bench Verified, LiveCodeBench, GPQA Diamond, HLE) are independently verified, but its agentic pitch is mostly DeepSeek's own math — three of its four agentic and long-horizon results (DeepSWE, Toolathlon-Verified, Agents' Last Exam) plus the tool-augmented HLE score are self-reported, with only Terminal-Bench independently run.

DeepSeek V4 Pro (0813) benchmarks and pricing, every number sourced: 11 tracked DeepSeek V4 Pro (0813) benchmark scores (7 independently run, 4 still resting on a vendor’s own claim), priced at $1.32 per million input tokens and $3.96 per million output.

DeepSeek V4 Pro (0813)’s 11 benchmark scores on this page were each verified against their sources between 2026-08-13 and 2026-08-24.

Released
2026-08-13
License
open-weights
Context window
1M tokens
Knowledge cutoff
Not disclosed
Parameters
Not disclosed
Architecture
Not disclosed

DeepSeek V4 Pro (0813)’s verified record

DeepSeek V4 Pro (0813)’s featured comparison is Claude Opus 4.8: 1 trail, 1 tie, and 8 not callable. DeepSeek V4 Pro (0813) is priced at $1.32/$3.96 per 1M tokens (in/out) vs Claude Opus 4.8’s $5.00/$25.00. Full verdict →

Against the 229 head-to-head comparisons DeepSeek V4 Pro (0813) shares with other tracked models: 42 real gaps, 36 inside the noise band, and 151 we will not call.

A gap counts for DeepSeek V4 Pro (0813) only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — DeepSeek V4 Pro (0813) trails on 33 of them.

HLE · no tools Reasoning

±2 is noise
Behind
Qwen3.8-Max −2.1 · Grok 4.7 −2.1 · Step 5 Preview −5.5 · Kimi K3 −5.9 · Gemini 3.1 Pro Preview −6.0 · Gemini 3.8 Flash −6.8 · GPT-6 Sol −6.9 · Muse Spark 1.3 −7.7 · MiMo-V2.6-Pro −8.4 · GPT-5.6 Sol −8.5 · GPT-6.1 Sol −11.9 · GPT-6 Astra −13.7 · Claude Opus 5 −13.9 · Claude Sonnet 5.5 −14.0 · Gemini 4 Argon −16.1 · Claude Fable 5.1 −18.1 · Claude Opus 5.5 −20.4 · 4 superseded: Muse Spark 1.2 −4.5 · Gemini 3.7 Flash −6.9 · Claude Opus 4.8 −7.7 · Claude Fable 5 −14.5
Ahead
GPT-6 Luna +2.5 · Qwen3.8-Flash-Next +3.0 · 1 superseded: DeepSeek V4 Flash (0731) +2.4
Tie
6 models within ±2
Unverified
2 models — vendor-reported on one side

LiveBench Composite score across 7 domains

±2.7 is noise
Behind
GPT-5.6 Sol −3.6 · Muse Spark 1.3 −4.2 · GPT-6.1 Sol −4.2 · Claude Opus 5.5 −5.8 · Claude Fable 5.1 −6.0 · 1 superseded: Claude Fable 5 −5.6
Ahead
GPT-5.6 Luna +3.8 · GPT-6 Luna +5.4 · GLM-5.3-Flash +5.8 · MiniMax M3 +10.1 · 2 superseded: DeepSeek V4 Flash (0731) +3.2 · GLM-5.2 +4.2
Tie
14 models within ±2.7

ARC-AGI-2 · max Compositional visual reasoning

±9.2 is noise
Behind
Claude Fable 5.1 −28.7 · Claude Opus 5 −29.1 · GPT-5.6 Sol −31.2 · GPT-6 Astra −33.7 · 1 superseded: Claude Fable 5 −27.9
Tie
4 models within ±9.2

Terminal-Bench 2.1 Terminal ops

±10.6 is noise
Behind
Tie
12 models within ±10.6
Unverified
1 model — vendor-reported on one side
Setup-dependent
17 models — scored on a different harness

No verdict for DeepSeek V4 Pro (0813) anywhere on Agents' Last Exam, DeepSWE, HLE · with tools, Toolathlon-Verified (nothing independently confirmed on both sides); GPQA Diamond, LiveCodeBench, SWE-bench Verified (saturated).

  • None of DeepSeek V4 Pro (0813)’s coding comparisons are independently confirmed on both sides yet.

DeepSeek V4 Pro (0813) API pricing

$1.32 in / $3.96 out per 1M tokens — official pricing

What DeepSeek V4 Pro (0813) costs per job

DeepSeek V4 Pro (0813) cost for three reference workloads, computed from its list rates
WorkloadTokens in / outCost
One long chat turn100K / 10K$0.172
A codebase review1,000K / 100K$1.72
A day of agent work10,000K / 1,000K$17.16

Computed from DeepSeek V4 Pro (0813)’s list rates above — cache discounts, its off-peak, and batch tiers are not applied.

DeepSeek V4 Pro (0813) is one of 3 DeepSeek models tracked on this site, at these official list prices.

DeepSeek model pricing, official list rates
ModelIn / 1MOut / 1M
DeepSeek V4.1 Flash$0.30$1.20
DeepSeek V4 Pro (0813)$1.32$3.96
DeepSeek V4 Flash (0731)(superseded)$0.44$1.32

DeepSeek V4 Pro (0813) benchmark scores

DeepSeek V4 Pro (0813) benchmark scores, provenance, and source links
BenchmarkScore
HLE(no tools)[1]
Reasoning · ±2 is noise
HLE(with tools)
Reasoning · ±2 is noise
Terminal-Bench 2.1[2]
Terminal ops · ±10.6 is noise
DeepSWE[3]
Long-horizon coding · ±9.5 is noise
Toolathlon-Verified
Multi-tool chores · ±9.7 is noise
Agents' Last Exam[4]
Professional work · ±3.2 is noise
SWE-bench Verifiedsaturated[5]
Bug fixing — not ranked at any gap size
GPQA Diamondsaturated[6]
Expert science Q&A — not ranked at any gap size
LiveCodeBenchsaturated[7]
Contest coding — not ranked at any gap size
ARC-AGI-2(max)[8]
Compositional visual reasoning · ±9.2 is noise
LiveBench[9]
Composite score across 7 domains · ±2.7 is noise

Who ran these numbers: 7 of 11 independent — vals.ai (3), artificialanalysis.ai (2), arcprize.org (1), livebench.ai (1); vendor self-reported (4).

  1. HLE: AA's own run of 'DeepSeek V4 Pro 0813 (Reasoning, Max Effort)' — text-only, no tools. Replaces the launch chart's self-reported 42.7. AA's run of the older 0424 V4 Pro scored 37.5; don't conflate.
  2. Terminal-Bench 2.1: AA's own harness ('DeepSeek V4 Pro 0813 (max)'). The vendor's launch chart claims 87.9 — 9.2 points above this independent run, and tbench.ai's official board carries no DeepSeek entry at all (board max is 83.8, below the vendor claim). Cross-check (2026-10-01): vals.ai's archived Terminus 2 table lists 'DeepSeek V4 Pro 0813' at 54.68, 24 points under this AA row.
  3. DeepSWE: Independently corroborated by third-party launch coverage. Up from 12.8 in the April preview build — the steepest single-generation jump on this chart.
  4. Agents' Last Exam: Vendor launch-chart claim; the 0813 model is not on Snorkel's board (its only DeepSeek row predates 0813 and scored 12.4 Overall pass rate under OpenClaw). Kept self-reported until an independent 0813 run exists. Corrected 2026-08-19: this row read 25.2, which is the adjacent DeepSeek-V4-Flash-0731 column on the same launch table — the DeepSeek-V4-Pro-0813 cell reads 25.7. Note this is a vendor-harness figure sitting in a column otherwise pinned to Snorkel's Overall pass_rate_pct; it is not comparable like-for-like and is retained only because no independent 0813 run exists.
  5. SWE-bench Verified: vals.ai run of the exact 0813 model, rank 2/83 (482/500 resolved), bash-only harness (updated 2026-08-14).
  6. GPQA Diamond: vals.ai run of the 0813 model, rank 15/133 (updated 2026-08-15) — displayed tied with Claude Opus 4.8.
  7. LiveCodeBench: vals.ai run of the 0813 model, rank 11/138 (updated 2026-08-15). Distinct from base 'DeepSeek V4' at 87.48.
  8. ARC-AGI-2: DeepSeek V4 Pro (0813)'s official ARC-AGI-2 leaderboard row, dated 2026-08-13 on arcprize.org. The highest of four populated tiers (Max/High/Low/None) for this row.
  9. LiveBench: Board row "DeepSeek V4 Pro 0813" on the 2026-06-25 LiveBench release.

Notes on the record

Service status rechecked 2026-09-17: DeepSeek's September 10 changelog (https://api-docs.deepseek.com/updates/) says V4 Pro API service continues after September 14 with billing unchanged, and promises further notice of changes. This site has therefore removed its superseded label and withdrawn the earlier claim that Pro requests would automatically route to V4.1 Flash. This is verification of the published service policy, not a live inference test.

Peak cache-miss rate shown; off-peak is exactly half — see the FAQ below for the exact schedule. Both rates are the undiscounted peak/off-peak basis every other model on this site uses (no vendor's time-of-day or batch discount is applied anywhere). MIT-licensed weights on Hugging Face. Confirmed against DeepSeek's own pricing page (api-docs.deepseek.com/quick_start/pricing) as of 2026-08-20: the peak/off-peak rates match exactly, cache-hit pricing is $0.044/1M tokens at peak and $0.022/1M off-peak, and no future price-change date is published — DeepSeek's terms only say prices "may vary" without a schedule.

Context window is 1M tokens with a 384K-token max output. The 0813 build is DeepSeek's general-availability release of V4 Pro, succeeding the preview version shipped in April 2026, and adds a DSpark speculative-decoding module for faster inference; the Hugging Face model card (huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813) also documents three selectable reasoning-effort levels — low, high, and max.

Corrected 2026-08-19: the listed price previously used DeepSeek's off-peak rate while V4 Flash used its peak rate, understating V4 Pro's price relative to Flash (1.5x instead of the true 3x) — both are now shown on the same undiscounted peak basis.

Compare with

FAQ

Has DeepSeek V4 Pro (0813) been independently benchmarked?

Partly. Seven of the eleven scores tracked for DeepSeek V4 Pro (0813) come from outside evaluators observed mid-August 2026 or later — Humanity's Last Exam without tools, Terminal-Bench 2.1, SWE-bench Verified, GPQA Diamond and LiveCodeBench (both via Vals.ai), and ARC-AGI-2 and LiveBench. The other four are DeepSeek's own self-reported runs: the tool-augmented HLE score, DeepSWE, Toolathlon-Verified, and Agents' Last Exam. The split falls along a fairly clean line — the standard coding and knowledge benchmarks have outside confirmation, while the harder tool-using and multi-step agentic evaluations are the ones DeepSeek reports on its own.

Why does DeepSeek V4 Pro (0813)'s price change during the day?

DeepSeek bills V4 Pro (0813) on a peak/off-peak schedule rather than a flat rate. Peak pricing — $1.32 per million input tokens and $3.96 per million output tokens on a cache miss — applies during 01:00-04:00 and 06:00-10:00 UTC. The remaining 17 of 24 hours are off-peak, at exactly half those rates ($0.66/$1.98). This site lists DeepSeek V4 Pro (0813) at the peak rate, the same undiscounted basis used for every other model on the site, since no vendor's time-of-day or batch discount is applied anywhere else. Confirmed against DeepSeek's own pricing page as of 2026-08-20, which does not publish a future change date for these rates.

Is DeepSeek V4 Pro (0813) open source?

Yes. DeepSeek publishes the V4 Pro (0813) weights on Hugging Face under the MIT license, which permits commercial use, modification, and redistribution without a copyleft requirement. That's separate from the hosted API pricing shown on this page, which only applies if you use DeepSeek's own endpoint rather than self-hosting the weights.

Is DeepSeek V4 Pro (0813) the same model as the preview DeepSeek shipped earlier in 2026?

No — 0813 is the general-availability release that succeeded the preview build DeepSeek shipped in April 2026 (per DeepSeek's own V4 announcement, deepseek.ai/blog/deepseek-v4-unveiled-1-6-trillion-parameters). The architecture and parameter count carry over from that preview — a mixture-of-experts model with roughly 1.6T total and around 49B active parameters — but the GA build adds a speculative-decoding module DeepSeek calls DSpark for faster inference, and is the version now served at DeepSeek's production API endpoint and reflected in the scores on this page.

Does DeepSeek V4 Pro (0813) let you control how much it reasons before answering?

Yes. Per DeepSeek's Hugging Face model card, V4 Pro (0813) exposes three reasoning-effort settings — low, high, and max — that trade latency for deliberation depth. DeepSeek's published benchmark results aren't broken out by effort level, so treat the scores on this page as representative rather than tied to one fixed setting.

Further reading

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.