Humanity's Last Exam
Expert-written questions across dozens of subjects, built to be hard enough that today's best models still fail most of them.
Items
2500
Trust grade
B
Status
current
How many points is real?
On HLE, a gap smaller than 2 points is treated as noise — with 2500 items, run-to-run variance alone can produce a difference that size. Anything bigger is worth taking seriously (subject to the caveats below).
In plain English: Grade B: Reliable, mostly vendor-reported, no major known issues — but still worth an independent check.
Trust caveats
Actively discriminating between frontier models.
2,500 questions from ~1,000 experts across 500+ institutions. Not saturated — top model as of Aug 2026 scores well under the ~88-90% range where benchmarks stop discriminating.
Current scores
| DeepSeek V4 Pro (0813)(with tools) | 60.0self-reported |
| Claude Opus 4.8(with tools) | 57.9self-reported |
| Qwen3.8-Max(with tools) | 56.2self-reported |
| Kimi K3(with tools) | 56.0self-reported |
| Claude Fable 5(no tools) | 55.5independent |
Source: https://arxiv.org/abs/2501.14249