DeepSeekPrevious version · DeepSeek V4.1 Flash
DeepSeek V4 Flash (0731)
All nine benchmark numbers on this page for DeepSeek V4 Flash (0731) are independent, third-party runs — none are DeepSeek's own claims. DeepSeek confirms the 2026-07-31 date in its own API changelog, but published no standalone launch post for it — the announcement is a changelog line, not the blog-and-benchmarks treatment its April preview got.
DeepSeek V4 Flash (0731) benchmarks and pricing, every number sourced: 9 tracked DeepSeek V4 Flash (0731) benchmark scores (9 independently run, 0 still resting on a vendor’s own claim), priced at $0.44 per million input tokens and $1.32 per million output.
DeepSeek V4 Flash (0731) architecture: Mixture-of-Experts; 284B total parameters (13B activated per token); 1M-token context window.
- Released
- 2026-07-31
- License
- open-weights
- Context window
- 1M tokens
- Knowledge cutoff
- Not disclosed
- Verified
- sources checked 2026-08-17–2026-10-03
- Parameters
- 284B (13B active)
- Architecture
- Mixture-of-Experts
DeepSeek V4 Flash (0731)’s verified record
DeepSeek V4 Flash (0731)’s most-compared rival is Kimi K3: 3 trails, 2 ties, and 4 not callable across their 9 shared comparisons. DeepSeek V4 Flash (0731) is priced at $0.44/$1.32 per 1M tokens (in/out) vs Kimi K3’s $3.00/$15.00.
Against the 207 head-to-head comparisons DeepSeek V4 Flash (0731) shares with other tracked models: 59 real gaps, 41 inside the noise band, and 107 we will not call.
A gap counts for DeepSeek V4 Flash (0731) only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — DeepSeek V4 Flash (0731) trails on 57 of them.
HLE · no tools Reasoning
±2 is noiseLiveBench Composite score across 7 domains
±2.7 is noiseDeepSWE Long-horizon coding
±9.5 is noiseARC-AGI-2 · max Compositional visual reasoning
±9.2 is noiseTerminal-Bench 2.1 Terminal ops
±10.6 is noiseToolathlon-Verified Multi-tool chores
±9.7 is noiseNo verdict for DeepSeek V4 Flash (0731) anywhere on GPQA Diamond, LiveCodeBench, SWE-bench Verified ( saturated).
DeepSeek V4 Flash (0731) API pricing
$0.44 in / $1.32 out per 1M tokens — official pricing
What DeepSeek V4 Flash (0731) costs per job
| Workload | Tokens in / out | Cost |
|---|---|---|
| One long chat turn | 100K / 10K | $0.057 |
| A codebase review | 1,000K / 100K | $0.572 |
| A day of agent work | 10,000K / 1,000K | $5.72 |
Computed from DeepSeek V4 Flash (0731)’s list rates above — cache discounts, its off-peak, and batch tiers are not applied.
DeepSeek V4 Flash (0731) is one of 3 DeepSeek models tracked on this site, at these official list prices.
| Model | In / 1M | Out / 1M |
|---|---|---|
| DeepSeek V4.1 Flash | $0.30 | $1.20 |
| DeepSeek V4 Pro (0813) | $1.32 | $3.96 |
| DeepSeek V4 Flash (0731)(previous version) | $0.44 | $1.32 |
DeepSeek V4 Flash (0731) benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise | |
| SWE-bench Verifiedsaturated[2] Bug fixing — not ranked at any gap size | |
| GPQA Diamondsaturated[3] Expert science Q&A — not ranked at any gap size | |
| LiveCodeBenchsaturated[4] Contest coding — not ranked at any gap size | |
| Toolathlon-Verified[5] Multi-tool chores · ±9.7 is noise | |
| DeepSWE[6] Long-horizon coding · ±9.5 is noise | |
| ARC-AGI-2(max)[7] Compositional visual reasoning · ±9.2 is noise | |
| LiveBench[8] Composite score across 7 domains · ±2.7 is noise | |
| Terminal-Bench 2.1[9] Terminal ops · ±10.6 is noise |
Who ran these numbers: 9 of 9 independent — vals.ai (3), artificialanalysis.ai (2), toolathlon.xyz (1), deepswe.datacurve.ai (1), arcprize.org (1), livebench.ai (1).
- HLE: AA's own run (Reasoning, Max Effort), from the AA model page — the model sits outside the HLE page's top-30 chart.
- SWE-bench Verified: vals.ai run — at $0.0099/test, the cheapest cost-per-test of any 85%+ scorer on the board (updated 2026-08-14).
- GPQA Diamond: vals.ai run (89.899), updated 2026-08-15.
- LiveCodeBench: vals.ai run (87.264), updated 2026-08-15. Corrected 2026-10-03: the value field read 87.3 (mis-rounded); the board displays 87.26.
- Toolathlon-Verified: 70.7±0.9 Pass@1 with Toolathlon's own 'Evaluated by us' badge (entry 2026-07-31). The pre-0731 V4 Flash scored 50.9 — a +19.8-point jump between minor versions, independently verified. Vendor self-reports 70.3, consistent.
- DeepSWE: 53%±4 on mini-swe-agent (board updated 2026-08-13); vendor self-reports 54.4, within the noise band.
- ARC-AGI-2: DeepSeek V4 Flash (0731)'s official ARC-AGI-2 leaderboard row, dated 2026-07-31 on arcprize.org. The highest of four populated tiers (Max/High/Low/None).
- LiveBench: Board row "DeepSeek V4 Flash 0731" on the 2026-06-25 LiveBench release.
- Terminal-Bench 2.1: AA's own run, row 'DeepSeek V4 Flash 0731 (max)' — underlying 0.786516853932584, board-displayed 78.7. Identical resolution to DeepSeek V4 Pro 0813's AA run (both 70/89 tasks). Fills the model's only gap in agent work: its TB 2.1 presence was previously AA-only via other models.
Notes on the record
Historical checkpoint: DeepSeek retired V4 Flash on September 10. Its changelog (https://api-docs.deepseek.com/updates/), rechecked 2026-09-17, says deepseek-v4-flash temporarily routes to V4.1 Flash. The scores and prices below describe the earlier 0731 checkpoint, not the model now served through that alias.
Historical pricing shown is the peak cache-miss headline rate (schedule effective 2026-08-16); off-peak is half ($0.22/$0.66) and cache hits cost ~$0.014/MTok at peak, ~$0.007/MTok off-peak. That 2026-08-16 16:00 UTC schedule replaced a flat $0.14/$0.28 that had held since launch (the prior schedule is not on DeepSeek's own docs, which publish only the current one — it is reconstructed from third-party pricing trackers, checked 2026-08-27), so that headline rate was an increase, not a launch price. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday; all other hours — weekends included — bill at the off-peak rate (source: api-docs.deepseek.com/quick_start/pricing, checked 2026-08-20 — this page also confirms the $0.44/$1.32 peak and $0.22/$0.66 off-peak figures held at that check, with no future change date posted).
304B total params (confirmed on the model's HuggingFace repo directly, huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731, checked 2026-08-20), MIT open weights on HuggingFace. The model card also lists three reasoning-effort settings (low/high/max) and recommends up to 384K output tokens for the high/max settings — 384K is also the max output length listed on the official pricing/docs page.
Release date, corrected 2026-08-28. This page used to say no announcement existed and that the date was inferred from version strings and timestamps. That was an error: DeepSeek's API changelog carries a dated 2026-07-31 entry announcing the public beta, mirrored in the zh-CN changelog, and the Models & Pricing table at that check mapped the API name deepseek-v4-flash to version DeepSeek-V4-Flash-0731. What is true is narrower — DeepSeek published no standalone launch post for this checkpoint, and its corporate news feed still tops out at the April preview, though its docs-site news section did get dedicated posts for the V4 Pro GA (2026-08-13) and Vision-Exp (2026-08-21).
Compare with
FAQ
Has DeepSeek V4 Flash (0731) been independently benchmarked?
Yes, fully. All nine scores tracked on this page for DeepSeek V4 Flash (0731) — HLE, SWE-bench Verified, GPQA Diamond, LiveCodeBench, Terminal-Bench 2.1, Toolathlon-Verified, DeepSWE, ARC-AGI-2, and LiveBench — are independent runs from third-party boards (Artificial Analysis, Vals.ai, Toolathlon, DeepSWE, ARC Prize, LiveBench), observed 2026-08-17 through 2026-10-03, not figures DeepSeek published itself. Terminal-Bench 2.1 joined on 2026-10-03, when Artificial Analysis's run (78.65) was added alongside the other boards. For example, the SWE-bench Verified score of 88.8 comes from Vals.ai, observed by this site 2026-08-17. That's 9-for-9 independent, with zero self-reported numbers on this page.
What does DeepSeek V4 Flash (0731) cost, and does the price change by time of day?
These are historical prices for the retired 0731 checkpoint, not current charges for its API alias. The August 16 schedule billed $0.44/$1.32 per million input/output tokens at peak and $0.22/$0.66 off-peak, with peak hours 01:00-04:00 and 06:00-10:00 UTC Monday through Friday. As of the September 17 changelog check, calls to deepseek-v4-flash temporarily route to V4.1 Flash; use that model's page for its pricing.
Is DeepSeek V4 Flash (0731) actually open source?
The weights are published on HuggingFace under the MIT license, which permits commercial use, modification, and redistribution without extra field-of-use restrictions. That remains separate from hosted API access: the 0731 checkpoint is retired from DeepSeek's API, and its old alias temporarily routes to V4.1 Flash, per the changelog checked 2026-09-17. Self-hosting the published weights preserves access to this checkpoint, with infrastructure costs rather than DeepSeek API charges.
Did DeepSeek ever officially announce DeepSeek V4 Flash (0731)?
Yes, but only in DeepSeek's API changelog, not with a launch post. A dated 2026-07-31 changelog entry announces the public beta and is mirrored in the zh-CN changelog; the Models & Pricing table checked in August mapped the API name deepseek-v4-flash to version DeepSeek-V4-Flash-0731. What DeepSeek never did is give this checkpoint a standalone blog post or press release — its corporate news feed at the August check still topped out at the April preview, even though the docs-site news section got dedicated posts for the V4 Pro GA and Vision-Exp releases that followed it.
Does DeepSeek V4 Flash (0731) support different reasoning modes?
Yes — per its HuggingFace model card, DeepSeek V4 Flash (0731) offers three reasoning-effort settings (low, high, max). DeepSeek's docs recommend allowing up to 384K output tokens when running the high or max settings, which is also the model's listed maximum output length on the official pricing page.
Does DeepSeek V4 Flash (0731) have any score that still ranks?
Yes. Terminal-Bench 2.1 at 78.65, an Artificial Analysis run recorded 2026-10-03, plus LiveBench at 74.2. Everything printed above those on its sheet — GPQA Diamond 89.9, SWE-bench Verified 88.8, LiveCodeBench 87.26 — sits on saturated boards.
Further reading
- GPQA Diamond leaderboard 2026 — DeepSeek V4 Flash (0731) is one of the 23 models it compares.
- GLM-5.3-Flash vs DeepSeek Flash vs Qwen3.8 — DeepSeek V4 Flash (0731) is one of the 3 models it compares.
- Models with 10M token context windows 2026 — DeepSeek V4 Flash (0731) is one of the 35 models it compares.