DeepSeek

DeepSeek V4 Flash (0731)

All six benchmark numbers on this page for DeepSeek V4 Flash (0731) are independent, third-party runs — none are DeepSeek's own claims. What's missing is any word from DeepSeek itself: no launch announcement was ever published, so the 2026-07-31 release date is reconstructed from version strings, Artificial Analysis metadata, HuggingFace timestamps, and the Toolathlon entry date rather than confirmed by the company.

DeepSeek V4 Flash (0731)’s 6 benchmark scores on this page were verified against their sources on or after 2026-08-17.

Released
2026-07-31
License
open-weights
Context window
1M tokens
Knowledge cutoff
Not disclosed

The verified record

Against the 72 head-to-head comparisons DeepSeek V4 Flash (0731) shares with other tracked models: 24 real gaps, 17 inside the noise band, and 31 we will not call.

A gap counts for DeepSeek V4 Flash (0731) only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — DeepSeek V4 Flash (0731) trails on 19 of them.

HLE · no toolsReasoning

±2 is noise
Behind
Claude Fable 5 16.9 · Claude Opus 5 16.3 · GPT-5.6 Sol 10.9 · Claude Opus 4.8 10.1 · Gemini 3.7 Flash 9.3 · Gemini 3.1 Pro Preview 8.4 · Kimi K3 8.3 · Qwen3.8-Max 4.4 · Grok 4.6 4.3 · GLM-5.3 3.7 · Claude Sonnet 5 2.7 · GLM-5.2 2.5 · DeepSeek V4 Pro (0813) 2.4

DeepSWELong-horizon coding

±9.5 is noise
Behind
Claude Opus 5 21.0 · GPT-5.6 Sol 20.0 · Claude Fable 5 17.0 · Kimi K3 16.0 · Grok 4.6 14.0 · Gemini 3.7 Flash 12.0
Tie
4 models within ±9.5
Unverified
2 models — vendor-reported on one side

LiveCodeBenchContest coding

±3.1 is noise
Ahead
GLM-5.2 +17.8 · GLM-5.3 +6.8 · Claude Sonnet 5 +4.9 · GPT-5.6 Sol +4.7
Tie
9 models within ±3.1

Toolathlon-VerifiedMulti-tool chores

±9.7 is noise
Ahead
GLM-5.2 +10.8
Tie
4 models within ±9.7
Unverified
3 models — vendor-reported on one side

No verdict for DeepSeek V4 Flash (0731) anywhere on GPQA Diamond, SWE-bench Verified (saturated).

DeepSeek V4 Flash (0731) API pricing

$0.44 in / $1.32 out per 1M tokens official pricing

DeepSeek V4 Flash (0731) is one of 2 DeepSeek models tracked on this site, at these official list prices.

DeepSeek model pricing, official list rates
ModelIn / 1MOut / 1M
DeepSeek V4 Pro (0813)$1.32$3.96
DeepSeek V4 Flash (0731)$0.44$1.32

DeepSeek V4 Flash (0731) benchmark scores

DeepSeek V4 Flash (0731) benchmark scores, provenance, and source links
BenchmarkScore
HLE(no tools)[1]
Reasoning · ±2 is noise
SWE-bench Verifiedsaturated[2]
Bug fixing — not ranked at any gap size
GPQA Diamondsaturated[3]
Expert science Q&A — not ranked at any gap size
LiveCodeBench[4]
Contest coding · ±3.1 is noise
Toolathlon-Verified[5]
Multi-tool chores · ±9.7 is noise
DeepSWE[6]
Long-horizon coding · ±9.5 is noise

Who ran these numbers: 6 of 6 independent — vals.ai (3), artificialanalysis.ai (1), toolathlon.xyz (1), deepswe.datacurve.ai (1).

  1. HLE: AA's own run (Reasoning, Max Effort), from the AA model page — the model sits outside the HLE page's top-30 chart.
  2. SWE-bench Verified: vals.ai run — at $0.0099/test, the cheapest cost-per-test of any 85%+ scorer on the board (updated 2026-08-14).
  3. GPQA Diamond: vals.ai run (89.899), updated 2026-08-15.
  4. LiveCodeBench: vals.ai run (87.264), updated 2026-08-15.
  5. Toolathlon-Verified: 70.7±0.9 Pass@1 with Toolathlon's own 'Evaluated by us' badge (entry 2026-07-31). The pre-0731 V4 Flash scored 50.9 — a +19.8-point jump between minor versions, independently verified. Vendor self-reports 70.3, consistent.
  6. DeepSWE: 53%±4 on mini-swe-agent (board updated 2026-08-13); vendor self-reports 54.4, within the error bar.

Notes on the record

Pricing shown is the peak cache-miss headline rate (schedule effective 2026-08-16); off-peak is half ($0.22/$0.66) and cache hits cost ~$0.014/MTok at peak, ~$0.007/MTok off-peak — the cheapest of the 14 models tracked on this site. Peak hours are 01:00-04:00 and 06:00-10:00 UTC; all other hours bill at the off-peak rate (source: api-docs.deepseek.com/quick_start/pricing, checked 2026-08-20 — this page also confirms the $0.44/$1.32 peak and $0.22/$0.66 off-peak figures still hold as of today, with no future change date posted).

304B total params (confirmed on the model's HuggingFace repo directly, huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731, checked 2026-08-20), MIT open weights on HuggingFace. The model card also lists three reasoning-effort settings (low/high/max) and recommends up to 384K output tokens for the high/max settings — 384K is also the max output length listed on the official pricing/docs page.

Shipped quietly — no announcement post exists; the release date rests on the version string, AA metadata, HF timestamps and the Toolathlon entry date, all agreeing on 2026-07-31.

Compare with

FAQ

Has DeepSeek V4 Flash (0731) been independently benchmarked?

Yes, fully. All six scores tracked on this page for DeepSeek V4 Flash (0731) — HLE, SWE-bench Verified, GPQA Diamond, LiveCodeBench, Toolathlon-Verified, and DeepSWE — are independent runs from third-party boards (Artificial Analysis, Vals.ai, Toolathlon, DeepSWE), observed 2026-08-17, not figures DeepSeek published itself. For example, the SWE-bench Verified score of 88.8 comes from Vals.ai, observed by this site 2026-08-17. That's 6-for-6 independent, with zero self-reported numbers on this page.

What does DeepSeek V4 Flash (0731) cost, and does the price change by time of day?

Yes. Per DeepSeek's pricing page (rate schedule effective 2026-08-16), DeepSeek V4 Flash (0731) bills $0.44 per million input tokens and $1.32 per million output tokens during peak hours — defined as 01:00-04:00 and 06:00-10:00 UTC — and exactly half that, $0.22/$0.66, the rest of the day. Cached input tokens are cheaper still: about $0.014/MTok at peak and $0.007/MTok off-peak. DeepSeek's docs note prices can be adjusted, but no future change date is posted as of 2026-08-20.

Is DeepSeek V4 Flash (0731) actually open source?

The weights are published on HuggingFace under the MIT license, which permits commercial use, modification, and redistribution without extra field-of-use restrictions. That's separate from API access, though: DeepSeek also serves DeepSeek V4 Flash (0731) through its own metered API at the pricing above, so the open license doesn't mean the hosted endpoint is free — self-hosting is the only way to use it at zero marginal API cost.

Why isn't there an official launch announcement for DeepSeek V4 Flash (0731)?

There isn't one on record. DeepSeek V4 Flash (0731) shipped without a blog post or press release from DeepSeek; its 2026-07-31 release date is inferred by cross-referencing the version string in the model name, Artificial Analysis's metadata, HuggingFace's upload timestamps, and the date it entered the Toolathlon leaderboard — all of which independently point to the same day.

Does DeepSeek V4 Flash (0731) support different reasoning modes?

Yes — per its HuggingFace model card, DeepSeek V4 Flash (0731) offers three reasoning-effort settings (low, high, max). DeepSeek's docs recommend allowing up to 384K output tokens when running the high or max settings, which is also the model's listed maximum output length on the official pricing page.

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.