Agents' Last Exam
Real, economically valuable professional work — video editing, CAD, manufacturing simulation — done end-to-end.
What does an Agents' Last Exam task look like?
One ALE task reproduces a single real deliverable that a working professional was actually paid to produce, staged inside a real OS sandbox with the relevant professional software already installed. The agent is given an instruction plus starting files (for example: a CAD/CAE project, a video project with source footage, or a 3D character mesh) and must operate the real application end-to-end, not a toy API, to produce a finished artifact. Grading then compares that artifact against a reference answer that stays hidden until after the agent finishes, using tolerance/exact-match/geometry checks rather than a human or AI judge eyeballing the result. This description follows the general task format and the illustrative task categories (a manufacturing simulation task, a video-compositing task, a character-animation task) that the ALE paper itself publishes to explain its evaluation design — it does not reproduce any task's actual instruction text or reference answer.
How many points on Agents' Last Exam is real?
On Agents' Last Exam, a gap smaller than 3.2 points is treated as noise — with 1000 items, run-to-run variance alone can produce a difference that size. The floor itself is derived, not asserted: 100 ÷ √1000 ≈ 3.2, two standard errors on a benchmark this size.
Grade A: Agents' Last Exam is designed to resist contamination (e.g. refreshed on a schedule). With 1000 items, trust close results here.
14 independent Agents' Last Exam scores — dots inside one dashed band are closer than Agents' Last Exam’s 3.2-point noise band, statistically indistinguishable. 6 vendor-reported Agents' Last Exam scores are not plotted — see the table below.
Can you trust Agents' Last Exam? Caveats
Agents' Last Exam is actively discriminating between frontier models — a gap of at least 3.2 points here is worth taking seriously, though it only counts as a confirmed real gap once both scores come from independent runs.
1,000+ tasks across 55 sub-fields from 300+ industry experts. Maintained by UC Berkeley's Center for Responsible, Decentralized Intelligence (RDI), led by Dawn Song, published as a 300+-author paper on 2026-06-03 (arXiv:2606.05405). Explicitly refreshed every ~6 months as a living benchmark to resist contamination — the best trust grade in this set for that reason. Snorkel AI's independently-run leaderboard, the source for most scores below, documents the refresh mechanic directly: "ALE uses rolling evaluation: every ~6 months a new public subset is published with fresh instances, while private tasks rotate in and retired public tasks rotate out, to limit benchmark leakage" (snorkel.ai/leaderboard/agents-last-exam). Berkeley RDI's own description of the grading discipline behind that leaderboard: "Every task is derived from a real project that a human expert previously completed and converted into a verifiable evaluation with objective grading. No vibes. No human judges. Fully reproducible" (rdi.berkeley.edu/blog/agents-last-exam).
Scoring mechanics
Each task deliverable is graded by a deterministic evaluator wherever one can be built — exact/hashed value matches, numeric fields with tolerances, geometric distances, or scripted world-state checks — run against a hidden reference that is staged only after the agent submits. The paper states ALE "deliberately avoids LLM-as-judge wherever a deterministic alternative exists" and rejects at QC any task whose only scoring path is "ask a model whether the result looks correct"; where an LLM (including a vision-LLM) judge is unavoidable, it is restricted to narrow, evidence-anchored yes/no probes rather than holistic "does this look right" grading — no specific judge model is named for these narrow probes in the published materials. Two distinct metrics are reported per model: a binary full pass rate (complete credit only, what our column tracks) and a separate continuous partial-credit mean score (0-1, credit for partially-correct work) — this is the pass_rate_pct vs score_pct distinction called out below. Because tasks are open-ended professional deliverables (a finished simulation, an edited video, a rigged character) rather than selected-response questions, there is no multiple-choice chance floor.
Metric pinned 2026-08-17
Our column stores the board's OVERALL-split pass_rate_pct for each model's best harness config (the board also publishes a partial-credit score_pct and a separate, harder ALE-CLI split — a terminal/Linux-only subset of the same task pool — do not mix; an earlier version of this dataset did).
Board update (checked 2026-10-01)
Snorkel's Overall table now has 43 rows. It adds Claude Opus 5.5 (38.2, rank 1), GPT-6 Astra (34.2), GPT-6 Sol (32.2), Muse Spark 1.3 (32.2) and GPT-6 Luna (25.0), all recorded here. Claude Opus 5's row changed: Snorkel's data notes say the upstream source re-published every Opus 5 figure on 2026-09-24, and this site's 27.0 (Max, score 49.0, est. cost $7,239, recorded 2026-08-20) is gone, replaced by 32.2 at High (score 55.9, est. cost $1,108) and 30.9 at Max; the eight other rows this site had recorded are unchanged. Rows run at several efforts now carry per-effort results, and the board shows each row's best effort by default. This site records that best-effort value, names the effort, and lists the per-effort figures in the row notes, since a best-effort pair can mix tiers (Claude Opus 5's best is High, Opus 5.5's is Max). The board has no row for nineteen tracked models: fourteen with no figure here at all, including Claude Fable 5.1, Claude Sonnet 5.5, Grok 4.7 and GLM-5.3, and five whose figure here is the vendor's own. The board's only DeepSeek V4 Pro row (12.4, OpenClaw) names no build; this site has treated it as a pre-0813 run since August and still excludes it.
Sources · 5
- Agents' Last Exam paper (arXiv:2606.05405, submitted 2026-06-03)
- ALE paper — grading methodology and ALE-CLI difficulty detail (arXiv HTML render)
- Snorkel AI — Agents' Last Exam independent leaderboard (rolling refresh policy, Overall/ALE-CLI splits, harness configs)
- Berkeley RDI — Agents' Last Exam blog post (maintainer, Dawn Song, "no human judges" quote)
- Berkeley RDI — official ALE-Benchmark GitHub repo (public reference-task format, hidden-reference grading, 150 public tasks)
Benchmarks related to Agents' Last Exam
- Toolathlon-Verified — Multi-tool chores, trust grade B
- Terminal-Bench 2.1 — Terminal ops, trust grade B
Is the current lead on Agents' Last Exam real?
Real gapClaude Opus 5.5 leads by 4.0 points on Agents' Last Exam — outside the noise band, worth acting on.
Written up in full, for models on this Agents' Last Exam board: Gemini 3.1 Pro Preview vs Kimi K3 · Claude Opus 4.8 vs Gemini 3.1 Pro Preview — every benchmark each pair shares, not just Agents' Last Exam.
Agents' Last Exam leaderboard: current scores
Agents' Last Exam’s 20 tracked scores were each verified against their source between 2026-08-17 and 2026-10-01.
| Model | Agents' Last Exam score |
|---|---|
| Claude Opus 5.5 | 38.2independent |
| GPT-6 Astra | 34.2independent |
| Claude Opus 5 | 32.2independent |
| GPT-6 Sol | 32.2independent |
| Muse Spark 1.3 | 32.2independent |
| GPT-5.6 Sol | 30.6independent |
| GPT-5.6 Luna | 30.3independent |
| Kimi K3 | 28.3independent |
| Claude Opus 4.8 | 27.0independent |
| Qwen3.8-Max | 27.0independent |
| Claude Fable 5 | 25.7independent |
| GPT-6 Luna | 25.0independent |
| GLM-5.2 | 20.4independent |
| Gemini 3.1 Pro Preview | 16.4independent |
| DeepSeek V4.1 Flash | 31.8self-reported |
| MiMo-V2.6-Pro | 31.6self-reported |
| MiMo-V2.6-Flash | 27.6self-reported |
| GLM-5.3-Flash | 26.3self-reported |
| DeepSeek V4 Pro (0813) | 25.7self-reported |
| Qwen3.8-Flash-Next | 24.3self-reported |
Agents' Last Exam FAQ
What does Agents' Last Exam measure?
Real, economically valuable professional work — video editing, CAD, manufacturing simulation — done end-to-end. The benchmark has 1000 items.
How big does a gap on Agents' Last Exam have to be to mean anything?
At least 3.2 points. Across 1000 items, a smaller gap sits inside the spread that repeat runs produce on their own, so we call it a tie — provided both scores are independent runs; if either side is a vendor claim, the row is labelled unverified instead.
Can you trust Agents' Last Exam scores?
We grade it A and list it as current — it still tells frontier models apart across 1000 items. A gap of at least 3.2 points clears that noise band, but it only counts as a real gap once both scores come from independent runs.
Source: https://arxiv.org/abs/2606.05405