Agents' Last Exam
Real, economically valuable professional work — video editing, CAD, manufacturing simulation — done end-to-end.
How many points is real?
On Agents' Last Exam, a gap smaller than 3.2 points is treated as noise — with 1000 items, run-to-run variance alone can produce a difference that size. Anything bigger is worth taking seriously (subject to the caveats below).
In plain English: Grade A: Designed to resist contamination (e.g. refreshed on a schedule). Trust close results here.
Trust caveats
Actively discriminating between frontier models.
1,000+ tasks across 55 sub-fields from 300+ industry experts. Explicitly refreshed every ~6 months as a living benchmark to resist contamination — the best trust grade in this set for that reason. Metric pinned 2026-08-17: our column stores the board's OVERALL-split pass_rate_pct for each model's best harness config (the board also publishes a partial-credit score_pct and an ALE-CLI split — do not mix; an earlier version of this dataset did).
Current scores
| GPT-5.6 Sol | 30.6independent |
| Kimi K3 | 28.3independent |
| Claude Opus 4.8 | 27.0independent |
| Qwen3.8-Max | 27.0independent |
| Claude Fable 5 | 25.7independent |
Source: https://arxiv.org/abs/2606.05405