Every model launch
comes with a chart.
Most of the gaps on it are noise. We read each AI model release and tell you which differences are real — and which you should ignore.
Humanity’s Last Exam · same two models
DeepSeek V4 Pro42.7
Claude Opus 4.849.8
Claude Opus 4.8 leads without tools. Same two models, same benchmark, opposite verdict when you flip one switch — the only thing that changed is the harness.
Numbers as published in DeepSeek’s launch chart (vendor-reported) — the table below swaps in independent runs where they exist.
Read the full verdict →Scores & pricing
Why no single score? →In plain English: There is no total score. We won’t collapse eight different benchmarks into one number that pretends a 0.5-point gap means something — see a head-to-head for which gaps here are real.