The Model Gap

Agents' Last Exam

Real, economically valuable professional work — video editing, CAD, manufacturing simulation — done end-to-end.

Items
1000
Trust grade
A
Status
current

How many points is real?

On Agents' Last Exam, a gap smaller than 3.2 points is treated as noise — with 1000 items, run-to-run variance alone can produce a difference that size. Anything bigger is worth taking seriously (subject to the caveats below).

In plain English: Grade A: Designed to resist contamination (e.g. refreshed on a schedule). Trust close results here.

Trust caveats

Actively discriminating between frontier models.

1,000+ tasks across 55 sub-fields from 300+ industry experts. Explicitly refreshed every ~6 months as a living benchmark to resist contamination — the best trust grade in this set for that reason. Metric pinned 2026-08-17: our column stores the board's OVERALL-split pass_rate_pct for each model's best harness config (the board also publishes a partial-credit score_pct and an ALE-CLI split — do not mix; an earlier version of this dataset did).

Current scores

GPT-5.6 Sol
Kimi K3
Claude Opus 4.8
Qwen3.8-Max
Claude Fable 5

Source: https://arxiv.org/abs/2606.05405