AI benchmark analysis and model verdicts
A single feed of every model verdict and every explainer we've published, in the order we published them.
Verdicts are head-to-head calls on a specific launch chart — DeepSeek V4 Pro vs. Claude Opus 4.8 is the first — read against the same real/tie/unverified rules as every table on this site. Blog posts step back from any one matchup to explain a pattern across the data, like how often vendor-reported scores drift from independent reruns. Both are held to the same standard: no number printed without a source, no claim that outruns what the source actually shows.
- Analysis2026-08-28
Models with 10M token context windows 2026
None of the 35 models we track ships a native 10M-token window. 24 run flat-rate to their max; 10 price-step at or above a threshold. Every claim sourced. - Analysis2026-08-27
GLM-5.3-Flash vs DeepSeek Flash vs Qwen3.8
GLM-5.3-Flash vs Qwen3.8-Flash-Next vs DeepSeek V4 Flash: independent-score counts ran 4/6, 0/4, 8/8 at launch — two days later the boards moved. - Analysis2026-08-24
GPQA Diamond leaderboard 2026
We graded our own GPQA Diamond leaderboard a D and stopped ranking it. Here is the data behind that call, and what changed across all 253 model pairs. - Verdict2026-08-17
DeepSeek V4 Pro benchmarks: real or noise?
Its launch chart shows it beating Claude Opus 4.8. About half of that chart deserves your attention. - Blog2026-08-17
Self-reported vs. independent AI benchmarks
When we counted, 29 of our 42 scores were vendor-reported. By day's end we'd swapped most for independent runs — one vendor claim was 5.6 points high.