The Model Gap

Every model launch
comes with a chart.

Most of the gaps on it are noise. We read each AI model release and tell you which differences are real — and which you should ignore.

Humanity’s Last Exam · same two models
DeepSeek V4 Pro42.7
Claude Opus 4.849.8

Claude Opus 4.8 leads without tools. Same two models, same benchmark, opposite verdict when you flip one switch — the only thing that changed is the harness.

Numbers as published in DeepSeek’s launch chart (vendor-reported) — the table below swaps in independent runs where they exist.

Read the full verdict →

Scores & pricing

Why no single score? →
Model🧠 Reasoning
no tools
💻 Terminal ops🛠️ Long-horizon coding💼 Professional workPrice in/out
DeepSeek V4 Pro (0813)
DeepSeek
$0.66 / $1.98
GPT-5.6 Sol
OpenAI
Not sourced yet$5.00 / $30.00
Gemini 3.1 Pro Preview
Google DeepMind
Not sourced yet$2.00 / $12.00
Kimi K3
Moonshot AI
$3.00 / $15.00
Qwen3.8-Max
Alibaba
$2.00 / $6.00
GLM-5.2
Zhipu AI (Z.ai)
$1.40 / $4.40
Claude Fable 5
Anthropic
$10.00 / $50.00
Claude Opus 5
Anthropic
Not sourced yetNot sourced yet$5.00 / $25.00
GLM-5.3
Zhipu AI (Z.ai)
Not sourced yetNot sourced yetNot sourced yetNot sourced yet
Claude Sonnet 5
Anthropic
Not sourced yet$2.00 / $10.00

In plain English: There is no total score. We won’t collapse eight different benchmarks into one number that pretends a 0.5-point gap means something — see a head-to-head for which gaps here are real.

Verdicts