DeepSeek V4 Pro benchmarks: real or noise?
August 17, 2026 · Verdict #1
From DeepSeek V4 Pro’s launch chart: the same two models on the same benchmark (Humanity’s Last Exam). Flip one test setting and the winner reverses — this is why you can’t read a launch chart at face value, and why this site exists.
Claude Opus 4.8 leads without tools. Same two models, same benchmark, opposite verdict when you flip one switch — the only thing that changed is the harness.
Numbers as published in DeepSeek’s launch chart (vendor-reported) — the table below swaps in independent runs where they exist.
Read the full verdict →On August 13, DeepSeek released V4 Pro with the now-standard launch artifact: a chart of agentic benchmarks where it trades wins with Claude Opus 4.8.
The chart is real data. It is also doing what launch charts always do: presenting every number at face value, as if they were all equally meaningful.
They are not. Here is the same chart, read honestly.
First, the provenance
Every number below comes from DeepSeek's own launch material. Vendor-reported, vendor-chosen benchmarks, vendor-chosen harness, no independent replication yet. That doesn't make the numbers false — it makes them unverified until third parties rerun them. Keep that discount rate in mind for everything that follows.
The headline flip: HLE, with and without tools
The most interesting row is Humanity's Last Exam, reported two ways: without tools and with tools.
| Setting | DeepSeek V4 Pro | Claude Opus 4.8 | Verdict |
|---|---|---|---|
| HLE, no tools | 42.7 | 49.8 | Opus by 7.1 |
| HLE, with tools | 60.0 | 57.9 | V4 Pro by 2.1 |
Same two models. Same benchmark. Opposite verdicts. The only thing that changed is the scaffolding.
Look at the deltas instead of the rankings: give V4 Pro tools and it gains +17.3 points. Give Opus the same tools and it gains +8.1. That asymmetry is the actual finding — DeepSeek clearly trained hard on tool use, and it shows.
But it also means "which model is smarter" is the wrong question here. Setup-dependent. If your workload gives models tools — agents, search, code execution — V4 Pro's number is the one that describes your world. If it doesn't, Opus is still ahead, by a real margin.
(A quiet footnote the headlines skipped: the rightmost column of DeepSeek's own chart, Fable 5, posts 53.3 and 63.0 on this same row — ahead in both settings. Launch charts pick their duels.)
The noise: Agents' Last Exam
V4 Pro scores 25.7. Opus 4.8 scores 25.7.
The same number, twice. Even without the noise argument there is nothing here to rank — and on benchmarks in this size class, run-to-run variance would swamp a gap far larger than zero anyway.
Tie. Anyone telling you otherwise is selling precision that does not exist.
(Correction, August 19, 2026: an earlier version of this section read V4 Pro's cell as 25.2 and called it a 0.5-point gap. 25.2 is the adjacent V4-Flash column on DeepSeek's own table — the V4 Pro cell reads 25.7. The verdict was a tie either way; the number was still wrong, so here is the note.)
The real story: DeepSWE
V4 Pro: 62.7. Opus 4.8: 58.0. A 4.7-point lead for DeepSeek — plausibly meaningful, pending replication.
But the solider number sits inside DeepSeek's own column: the previous release, V4 Pro Preview, scored 12.8 on this row. That is a jump of roughly fifty points in one generation, measured by the same vendor on the same harness both times — no cross-lab excuse available.
Real gap — not the duel with Opus, which needs a third party to confirm, but the trajectory. Whatever DeepSeek changed in its agentic software-engineering training between Preview and Pro, it worked.
For balance: on Toolathlon-Verified, Opus stays ahead 76.2 to 74.1 — a 2.1-point edge that is itself borderline noise. Call it a tie until someone reruns it.
The one thing to remember
DeepSeek V4 Pro genuinely caught up with frontier agents when given tools, and its generation-over-generation jump is real and steep. The "beats Opus 4.8" headline, though, lives inside the margin of error and inside one particular harness. If you run tool-using agents, add it to your test list. If you were about to pick a model off this chart's rankings — don't.
Updated August 18, 2026: the numbers above are DeepSeek's original launch chart, left unedited so you can see exactly what shipped on release day. Independent runs have since landed for several of these rows and mostly turn this matchup into ties — Opus's one row that holds up as a confirmed, independently-verified win is HLE without tools (48.7 vs 41.0). See Claude Opus 4.8 vs. DeepSeek V4 Pro for the current comparison.
All numbers above are from DeepSeek's launch chart (vendor-reported), as of August 13, 2026. The rule behind each label in bold above is written out in the methodology.
Why do vendor-reported numbers drift from independent runs in the first place? We measured it on this site's own data: self-reported vs. independent AI benchmarks.