Humanity's Last Exam
Expert-written questions across dozens of subjects, built to be hard enough that today's best models still fail most of them.
What does an HLE task look like?
One HLE item is a single question written by a named subject-matter expert in one of dozens of academic fields — mathematics, physics, chemistry, biology, medicine, computer science, humanities, and more — and included specifically because the frontier models available when the dataset was built answered it wrong. About three-quarters of items are "exact-match": the model must output one specific short string (a number, a term, a formula) that an LLM judge checks against a fixed reference answer, tolerating equivalent formats. The remaining roughly one-quarter are multiple-choice, with several answer options rather than a fixed four-choice layout. A minority of items also pair the question text with an image the model has to interpret. HLE's own site (lastexam.ai) has published real sample questions publicly; one such public sample is a graduate-level biology question that asks for a single precise number as its answer, illustrating the level of domain specificity involved — its exact wording is not reproduced here.
How many points on HLE is real?
On HLE, a gap smaller than 2 points is treated as noise — with 2500 items, run-to-run variance alone can produce a difference that size. The floor itself is derived, not asserted: 100 ÷ √2500 ≈ 2, two standard errors on a benchmark this size.
Grade B: No major known issues on HLE, and the results here hold up — but with 2500 items, close calls still deserve an independent check.
30 independent HLE scores — dots inside one dashed band are closer than HLE’s 2-point noise band, statistically indistinguishable. 2 vendor-reported HLE scores are not plotted — see the table below.
Can you trust HLE? Caveats
HLE is actively discriminating between frontier models — a gap of at least 2 points here is worth taking seriously, though it only counts as a confirmed real gap once both scores come from independent runs.
2,500 questions from ~1,000 experts across 500+ institutions in 50 countries, created by the Center for AI Safety and Scale AI and later published in Nature (Jan 2026, DOI 10.1038/s41586-025-09962-4). Not saturated — top model as of Aug 2026 scores well under the ~88-90% range where benchmarks stop discriminating.
Scoring mechanics
Per the original paper, about 76% of items are exact-match short answers, graded for equivalence by an o3-mini LLM judge (it accepts formatting variants such as decimals vs. fractions); the remaining ~24% are multiple-choice, with a variable number of options rather than a fixed 4-choice format. There is no single chance floor for the benchmark as a whole — most of it is free-response with an effectively 0% guessing floor, and the multiple-choice minority doesn't have a uniform option count to compute one floor from. HLE separately tracks RMS calibration error (stated confidence 0-100% vs. actual accuracy, using the method from Hendrycks et al. 2022) alongside plain accuracy; this site tracks only the accuracy figure. The paper also discloses a private held-out question set kept off the public release specifically to help detect training-data contamination and overfitting.
Known answer-key problem, now reasonably well-documented
FutureHouse published an audit on 2025-07-23 (updated 2025-09-16) finding 29%±3.7% (95% CI) of the 321 text-only chemistry/biology answers were directly contradicted by peer-reviewed literature, using its Crow literature-search agent with expert cross-checking. HLE's own organizers (Center for AI Safety/Scale AI) responded with an independent three-expert review of a bio/chem subset, replicating a smaller but real ~18% problem rate (reviewers disagreed on 25% of items reviewed) and launched a rolling revision process to strip and replace flawed questions. A separate, independent HLE-Verified paper (arXiv 2602.13964, Qwen Team/Alibaba, first posted Feb 2026 — not affiliated with HLE's original creators) then re-audited the full set: 668 items verified clean, 1,143 revised, 689 left flagged as an unresolved uncertain set; re-scoring eight frontier models on the cleaned version moved average accuracy by 7-10 points overall and by 30-40 points specifically on the items that had been flawed — concrete evidence the answer-key noise is large enough to matter for close model-to-model comparisons. Graded B not A for this reason: the noise is concentrated in two subjects rather than site-wide, and it inflates and deflates different models unpredictably rather than in one direction — but treat single-digit gaps on chem/bio-heavy runs with extra caution.
This benchmark is this site's founding example of why variant labels matter. On the independently verified no_tools variant, Claude Opus 4.8 leads DeepSeek V4 Pro by 7.7 points (48.7 vs. 41.0). On the with_tools variant, DeepSeek V4 Pro reports 60.0 vs. Claude Opus 4.8's 57.9 — a lead running the other direction. Neither with_tools number is independently verified: every with_tools HLE score tracked on this site is vendor-self-reported, and in this specific case Opus's with_tools figure is sourced from DeepSeek's own published comparison chart rather than from Anthropic's materials. Read as one ranking, HLE flips depending on tool access; read correctly, only the no_tools half of that flip is confirmed by an independent board — the with_tools half is two vendors' unverified numbers pointing the other way.
Sources · 7
- Humanity's Last Exam paper (arXiv 2501.14249)
- HLE published in Nature, Jan 2026
- FutureHouse: "About 30% of Humanity's Last Exam Answers are Wrong" (2025-07-23, updated 2025-09-16)
- The Decoder coverage of the FutureHouse audit
- HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam (arXiv 2602.13964)
- lastexam.ai (official HLE site, public sample questions)
- Artificial Analysis HLE leaderboard (independent no_tools scoring board referenced in given score data)
Benchmarks related to HLE
- GPQA Diamond — Expert science Q&A, trust grade D
- ARC-AGI-2 — Compositional visual reasoning, trust grade B
Is the current lead on HLE real?
Real gapClaude Opus 5.5 leads by 2.3 points on Humanity's Last Exam — outside the noise band, worth acting on.
Computed on the no tools variant, the same one charted above — other variants of HLE can rank differently and are shown separately in the table below.
Written up in full, for models on this HLE board: Claude Opus 5 vs Claude Sonnet 5 · Claude Opus 4.8 vs DeepSeek V4 Pro (0813) · Claude Opus 4.8 vs Gemini 3.1 Pro Preview — every benchmark each pair shares, not just HLE.
HLE leaderboard: current scores
HLE’s 43 tracked scores were each verified against their source between 2026-08-13 and 2026-09-30.
| Model | HLE score |
|---|---|
| Claude Sonnet 5.5 | 64.5self-reported |
| DeepSeek V4.1 Flash | 63.9self-reported |
| GLM-5.3 | 62.5self-reported |
| DeepSeek V4 Pro (0813) | 60.0self-reported |
| Step 5 Preview | 59.4self-reported |
| Claude Opus 4.8 | 57.9self-reported |
| Qwen3.8-Max | 56.2self-reported |
| Kimi K3 | 56.0self-reported |
| Tencent Hy4 preview | 55.4self-reported |
| GLM-5.2 | 54.7self-reported |
| Gemini 3.1 Pro Preview | 51.4self-reported |
HLE FAQ
What does HLE measure?
Expert-written questions across dozens of subjects, built to be hard enough that today's best models still fail most of them. The benchmark has 2500 items.
How big does a gap on HLE have to be to mean anything?
At least 2 points. Across 2500 items, a smaller gap sits inside the spread that repeat runs produce on their own, so we call it a tie — provided both scores are independent runs; if either side is a vendor claim, the row is labelled unverified instead.
Can you trust HLE scores?
We grade it B and list it as current — it still tells frontier models apart across 2500 items. A gap of at least 2 points clears that noise band, but it only counts as a real gap once both scores come from independent runs.
How many questions are in Humanity's Last Exam?
2,500, written by roughly 1,000 experts across more than 500 institutions in 50 countries. About 76% are exact-match short answers graded by an LLM judge for equivalence, and the remaining 24% are multiple choice with a varying number of options — so there is no single guessing floor for the exam as a whole. A separate private held-out set exists, kept off the public release to detect contamination.
Who made Humanity's Last Exam, and is it peer reviewed?
It was built by the Center for AI Safety and Scale AI, first posted as arXiv:2501.14249, and later published in Nature in January 2026. Worth being precise about what that review covers: the benchmark's construction and methodology, not the correctness of every individual answer key. Those are a separate and well-documented problem, below.
Are the answers in Humanity's Last Exam actually correct?
A meaningful share are not, and this is documented rather than rumoured. FutureHouse audited the 321 text-only chemistry and biology items and found 29% (±3.7%) directly contradicted by peer-reviewed literature. HLE's own organisers ran an independent three-expert review, replicated a smaller but real ~18% problem rate, and started a rolling revision process. A team at Alibaba then re-audited the full set: 668 items clean, 1,143 revised, 689 still flagged. Re-scoring eight frontier models on the cleaned version moved accuracy by 7-10 points overall and 30-40 points on the flawed items. The 2-point noise floor this site applies to HLE still stands; what this adds is that chemistry- and biology-heavy runs deserve extra caution on top of it, because the error is concentrated in those two subjects and does not push every model the same way.
Source: https://arxiv.org/abs/2501.14249