The Model Gap

Humanity's Last Exam

Expert-written questions across dozens of subjects, built to be hard enough that today's best models still fail most of them.

Items
2500
Trust grade
B
Status
current

How many points is real?

On HLE, a gap smaller than 2 points is treated as noise — with 2500 items, run-to-run variance alone can produce a difference that size. Anything bigger is worth taking seriously (subject to the caveats below).

In plain English: Grade B: Reliable, mostly vendor-reported, no major known issues — but still worth an independent check.

Trust caveats

Actively discriminating between frontier models.

2,500 questions from ~1,000 experts across 500+ institutions. Not saturated — top model as of Aug 2026 scores well under the ~88-90% range where benchmarks stop discriminating.

Current scores

DeepSeek V4 Pro (0813)(with tools)
Claude Opus 4.8(with tools)
Qwen3.8-Max(with tools)
Kimi K3(with tools)
Claude Fable 5(no tools)

Source: https://arxiv.org/abs/2501.14249