DeepSeek

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash launched on September 10 with a vendor claim of outperforming V4 Pro, but all six benchmark scores recorded here are DeepSeek's own numbers. The launch-day check found no independent row on the eleven tracked benchmarks. V4 Flash is retired; V4 Pro API service remains available with unchanged billing, per the changelog rechecked September 17.

DeepSeek V4.1 Flash benchmarks and pricing, every number sourced: 6 tracked DeepSeek V4.1 Flash benchmark scores (0 independently run, 6 still resting on a vendor’s own claim), priced at $0.30 per million input tokens and $1.20 per million output.

DeepSeek V4.1 Flash is Mixture-of-Experts (Causal Encoder-Decoder) with 552B total parameters (8B prefill / 16B decode activated per token) and a 1M-token context window.

DeepSeek V4.1 Flash’s 6 benchmark scores on this page were verified against their source on 2026-09-10.

Version history: succeeded DeepSeek V4 Flash (0731) (2026-07-31).

Released
2026-09-10
License
open-weights
Context window
1M tokens
Knowledge cutoff
Not disclosed
Parameters
552B (8B prefill / 16B decode active)
Architecture
Mixture-of-Experts (Causal Encoder-Decoder)

DeepSeek V4.1 Flash’s verified record

DeepSeek V4.1 Flash’s most-compared rival is DeepSeek V4 Pro (0813): 6 not callable across their 6 shared comparisons. DeepSeek V4.1 Flash is priced at $0.30/$1.20 per 1M tokens (in/out) vs DeepSeek V4 Pro (0813)’s $1.32/$3.96.

Against the 124 head-to-head comparisons DeepSeek V4.1 Flash shares with other tracked models: 0 real gaps, 0 inside the noise band, and 124 we will not call.

A gap counts for DeepSeek V4.1 Flash only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — DeepSeek V4.1 Flash trails on 0 of them.

No verdict for DeepSeek V4.1 Flash anywhere on Agents' Last Exam, DeepSWE, HLE · no tools, HLE · with tools, Terminal-Bench 2.1 (nothing independently confirmed on both sides); GPQA Diamond (saturated).

  • None of DeepSeek V4.1 Flash’s reasoning comparisons are independently confirmed on both sides yet.
  • None of DeepSeek V4.1 Flash’s coding comparisons are independently confirmed on both sides yet.
  • None of DeepSeek V4.1 Flash’s agentic comparisons are independently confirmed on both sides yet.

What changed from DeepSeek V4 Flash (0731) to DeepSeek V4.1 Flash

The 3 benchmarks both models have been scored on, using the same variant each time. A raw DeepSeek V4.1 Flash gain is not a real gain until it clears that benchmark’s own noise band, so each row below carries the verdict and not just the arithmetic.

DeepSeek V4.1 Flash versus DeepSeek V4 Flash (0731), per-benchmark change and whether it clears the noise band
BenchmarkDeepSeek V4 Flash (0731)DeepSeek V4.1 FlashChangeVerdict
HLE38.639.1+0.5Unverified
DeepSWE5374.2+21.2Unverified
GPQA Diamond89.990.9+1.0Tainted

DeepSeek V4.1 Flash API pricing

$0.30 in / $1.20 out per 1M tokens official pricing

What DeepSeek V4.1 Flash costs per job

DeepSeek V4.1 Flash cost for three reference workloads, computed from its list rates
WorkloadTokens in / outCost
One long chat turn100K / 10K$0.042
A codebase review1,000K / 100K$0.420
A day of agent work10,000K / 1,000K$4.20

Computed from DeepSeek V4.1 Flash’s list rates above — cache discounts, its off-peak, and batch tiers are not applied.

DeepSeek V4.1 Flash is one of 3 DeepSeek models tracked on this site, at these official list prices.

DeepSeek model pricing, official list rates
ModelIn / 1MOut / 1M
DeepSeek V4.1 Flash$0.30$1.20
DeepSeek V4 Pro (0813)$1.32$3.96
DeepSeek V4 Flash (0731)(superseded)$0.44$1.32

DeepSeek V4.1 Flash benchmark scores

DeepSeek V4.1 Flash benchmark scores, provenance, and source links
BenchmarkScore
GPQA Diamondsaturated[1]
Expert science Q&A — not ranked at any gap size
HLE(no tools)[2]
Reasoning · ±2 is noise
HLE(with tools)[3]
Reasoning · ±2 is noise
Terminal-Bench 2.1[4]
Terminal ops · ±10.6 is noise
DeepSWE[5]
Long-horizon coding · ±9.5 is noise
Agents' Last Exam[6]
Professional work · ±3.2 is noise

Who ran these numbers: 0 of 6 independent; vendor self-reported (6).

  1. GPQA Diamond: Vendor launch-chart claim (Max reasoning effort). GPQA Diamond is graded saturated on this site, so the number carries little weight either way; kept self-reported until an independent run exists — vals.ai's live board (checked 2026-09-10) does not yet list this model.
  2. HLE: DeepSeek's own comparison table lists two no-tools HLE figures for this model: 39.1 on the text-only subset (footnoted †) and 36.8 on the full, multimodal question set. This row tracks the text-only figure, matching this site's no_tools convention for every other model (sourced from Artificial Analysis's text-only methodology elsewhere on this site) — the vendor's own primary headline number, 36.8, is the wider (harder) multimodal set and is not the comparable one. No independent run exists yet.
  3. HLE: Vendor launch-chart claim (Max reasoning effort, tool use enabled). No independent run exists yet.
  4. Terminal-Bench 2.1: DeepSeek's own "DSH Minimal" harness (DeepSeek Harness, Minimal mode) — not tbench.ai's Terminus 2 nor Artificial Analysis's harness, neither of which lists this model yet (both checked live 2026-09-10). DeepSeek's own scaffold-comparison table shows this is the highest of eight scaffolds it tested: Claude Code scored 88.0, Codex 84.1, OpenCode 85.0, Pi 86.1, mini-SWE 90.3, DSH Standard 85.8, DSH PTC 85.8 on the same model.
  5. DeepSWE: Vendor launch-chart claim (Max reasoning effort), run on the mini-SWE-agent harness — the same harness the independent deepswe.datacurve.ai board itself uses, so this figure is closer to reproducible than most vendor claims on this site, but it is still DeepSeek's own run, not deepswe.datacurve.ai's. That board (checked live 2026-09-10) does not list this model yet.
  6. Agents' Last Exam: Vendor launch-chart claim, per DeepSeek's own methodology run on ALE's official scaffold. Not from Snorkel AI's independent leaderboard (the source this site uses for this benchmark), whose only DeepSeek row — a pre-0813 DeepSeek V4 Pro checkpoint that scored 12.4 Overall pass rate under OpenClaw — is a different model on a different harness. Kept self-reported until an independent run exists.

Notes on the record

DeepSeek released V4.1 Flash on 2026-09-10. Service status rechecked against its September 10 changelog (https://api-docs.deepseek.com/updates/) on 2026-09-17: V4 Flash and V4 Flash Vision Exp are retired, with deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily routed to V4.1 Flash. V4 Pro API service continues beyond September 14 with billing unchanged. The earlier claim on this page that Pro requests would redirect at 04:00 UTC on that date is withdrawn; no inference request was made to test routing. Pricing, confirmed on DeepSeek's own USD page (checked 2026-09-10): $0.15/$0.60 per million input/output tokens off-peak, cache-miss basis, $0.30/$1.20 at peak (01:00–04:00 and 06:00–10:00 UTC, Mon–Fri). Context window 1M tokens, max output 384K, unchanged from its predecessors.

Per DeepSeek's technical report (huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash, checked 2026-09-10), V4.1 Flash is a natively multimodal Mixture-of-Experts model with a Causal Encoder-Decoder architecture: 552B backbone parameters, activating 8B per token at prefill and 16B at decode, plus a separate 196B-parameter Engram memory. Hugging Face's own weight-file count reads 763B total — the technical report does not reconcile that figure against the 552B backbone / 196B memory breakdown it states in prose, so the extra roughly-15B is unexplained by the card itself, not necessarily just the vision encoder. MIT-licensed on Hugging Face, the same open posture as V4 Pro and V4 Flash. Reasoning effort is continuous (1–100) per the technical report, but the live API still exposes only the same three discrete tiers as its predecessors — low, high, max.

All six benchmark scores tracked for it here are DeepSeek's own self-reported numbers, identical across the Hugging Face card and the 2026-09-10 changelog. None of the eleven benchmarks this site tracks has an independent leaderboard entry for this model yet — vals.ai, tbench.ai, Artificial Analysis, deepswe.datacurve.ai, toolathlon.xyz, arcprize.org, livebench.ai, matharena.ai and Snorkel AI were all checked live on 2026-09-10 and none list it, consistent with a same-day release.

Compare with

FAQ

Has DeepSeek V4.1 Flash been independently benchmarked yet?

No — not on any of the eleven benchmarks this site tracks. DeepSeek V4.1 Flash launched 2026-09-10, and as of that date none of vals.ai, tbench.ai, Artificial Analysis, deepswe.datacurve.ai, toolathlon.xyz, arcprize.org, livebench.ai, matharena.ai, or Snorkel AI had added it to their boards. The six scores shown on this page — GPQA Diamond, both variants of HLE, Terminal-Bench 2.1, DeepSWE, and Agents' Last Exam — are all DeepSeek's own self-reported numbers from its technical report and changelog.

Does DeepSeek V4.1 Flash replace DeepSeek V4 Pro and V4 Flash?

It replaces V4 Flash, but V4 Pro API service remains available. DeepSeek's changelog, rechecked 2026-09-17, says the old Flash and Flash Vision Exp API names temporarily route to V4.1 Flash, while Pro continues after September 14 with billing unchanged. The previous answer here incorrectly presented a Pro redirect as settled. The V4 Flash page retains its historical scores and superseded label; the V4 Pro page is current.

What does DeepSeek V4.1 Flash cost, and does the price change by time of day?

Yes — like earlier V4 models, DeepSeek bills V4.1 Flash on a peak/off-peak schedule. Off-peak (most of the day) runs $0.15 per million input tokens and $0.60 per million output tokens on a cache miss; peak hours — 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday — double both to $0.30/$1.20. Cached input tokens are far cheaper: $0.003 off-peak, $0.006 at peak. That undercuts the earlier V4 models' off-peak input rates — $0.22 for V4 Flash, $0.66 for V4 Pro — confirmed on DeepSeek's own USD pricing page, checked 2026-09-10.

Is DeepSeek V4.1 Flash open source?

Yes. DeepSeek publishes the V4.1 Flash weights on Hugging Face under the MIT license, the same posture as V4 Pro and V4 Flash — permitting commercial use, modification, and redistribution without a copyleft requirement. That's separate from the hosted API pricing above, which only applies if you call DeepSeek's own endpoint rather than self-hosting; Hugging Face lists deployment instructions for vLLM, SGLang, Transformers, and llama.cpp, among others.

What's actually new about DeepSeek V4.1 Flash's architecture?

Per DeepSeek's own technical report, the headline change is KV-cache compression, not a bigger model: a new Causal Encoder-Decoder architecture activates only 8B parameters per token at prefill and 16B at decode (versus a flat 13B for V4 Flash), and combined with several attention and caching changes, DeepSeek claims roughly a 4x smaller KV-cache footprint per token than V4 Flash. The model is also natively multimodal — it processes images directly, which neither of the two earlier V4 models tracked here could; image input for the V4 generation shipped only as a separate experimental model, V4 Flash Vision Exp, retired the same day.

Further reading

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.