Analysis
A single feed of every model verdict and every explainer we've published, in the order we published them.
Verdicts are head-to-head calls on a specific launch chart — DeepSeek V4 Pro vs. Claude Opus 4.8 is the first — read against the same real/tie/unverified rules as every table on this site. Blog posts step back from any one matchup to explain a pattern across the data, like how often vendor-reported scores drift from independent reruns. Both are held to the same standard: no number printed without a source, no claim that outruns what the source actually shows.
- Verdict2026-08-17DeepSeek V4 Pro benchmarks: real or noise?Its launch chart shows it beating Claude Opus 4.8. About half of that chart deserves your attention.
- Blog2026-08-17Self-reported vs. independent AI benchmarksWhen we counted, 29 of our 42 scores were vendor-reported. By day's end we'd swapped most for independent runs — one vendor claim was 5.6 points high.