Agents' Last Exam

Real, economically valuable professional work — video editing, CAD, manufacturing simulation — done end-to-end.

Items
1000
Trust grade
A
Status
current

What does an Agents' Last Exam task look like?

One ALE task reproduces a single real deliverable that a working professional was actually paid to produce, staged inside a real OS sandbox with the relevant professional software already installed. The agent is given an instruction plus starting files (for example: a CAD/CAE project, a video project with source footage, or a 3D character mesh) and must operate the real application end-to-end, not a toy API, to produce a finished artifact. Grading then compares that artifact against a reference answer that stays hidden until after the agent finishes, using tolerance/exact-match/geometry checks rather than a human or AI judge eyeballing the result. This description follows the general task format and the illustrative task categories (a manufacturing simulation task, a video-compositing task, a character-animation task) that the ALE paper itself publishes to explain its evaluation design — it does not reproduce any task's actual instruction text or reference answer.

How many points on Agents' Last Exam is real?

On Agents' Last Exam, a gap smaller than 3.2 points is treated as noise — with 1000 items, run-to-run variance alone can produce a difference that size. The floor itself is derived, not asserted: 100 ÷ √1000 ≈ 3.2, two standard errors on a benchmark this size.

Grade A: Agents' Last Exam is designed to resist contamination (e.g. refreshed on a schedule). With 1000 items, trust close results here.

14 independent Agents' Last Exam scores — dots inside one dashed band are closer than Agents' Last Exam’s 3.2-point noise band, statistically indistinguishable. 6 vendor-reported Agents' Last Exam scores are not plotted — see the table below.

Can you trust Agents' Last Exam? Caveats

Agents' Last Exam is actively discriminating between frontier models — a gap of at least 3.2 points here is worth taking seriously, though it only counts as a confirmed real gap once both scores come from independent runs.

1,000+ tasks across 55 sub-fields from 300+ industry experts. Maintained by UC Berkeley's Center for Responsible, Decentralized Intelligence (RDI), led by Dawn Song, published as a 300+-author paper on 2026-06-03 (arXiv:2606.05405). Explicitly refreshed every ~6 months as a living benchmark to resist contamination — the best trust grade in this set for that reason. Snorkel AI's independently-run leaderboard, the source for most scores below, documents the refresh mechanic directly: "ALE uses rolling evaluation: every ~6 months a new public subset is published with fresh instances, while private tasks rotate in and retired public tasks rotate out, to limit benchmark leakage" (snorkel.ai/leaderboard/agents-last-exam). Berkeley RDI's own description of the grading discipline behind that leaderboard: "Every task is derived from a real project that a human expert previously completed and converted into a verifiable evaluation with objective grading. No vibes. No human judges. Fully reproducible" (rdi.berkeley.edu/blog/agents-last-exam).

Scoring mechanics

Each task deliverable is graded by a deterministic evaluator wherever one can be built — exact/hashed value matches, numeric fields with tolerances, geometric distances, or scripted world-state checks — run against a hidden reference that is staged only after the agent submits. The paper states ALE "deliberately avoids LLM-as-judge wherever a deterministic alternative exists" and rejects at QC any task whose only scoring path is "ask a model whether the result looks correct"; where an LLM (including a vision-LLM) judge is unavoidable, it is restricted to narrow, evidence-anchored yes/no probes rather than holistic "does this look right" grading — no specific judge model is named for these narrow probes in the published materials. Two distinct metrics are reported per model: a binary full pass rate (complete credit only, what our column tracks) and a separate continuous partial-credit mean score (0-1, credit for partially-correct work) — this is the pass_rate_pct vs score_pct distinction called out below. Because tasks are open-ended professional deliverables (a finished simulation, an edited video, a rigged character) rather than selected-response questions, there is no multiple-choice chance floor.

Metric pinned 2026-08-17

Our column stores the board's OVERALL-split pass_rate_pct for each model's best harness config (the board also publishes a partial-credit score_pct and a separate, harder ALE-CLI split — a terminal/Linux-only subset of the same task pool — do not mix; an earlier version of this dataset did).

Board update (checked 2026-10-01)

Snorkel's Overall table now has 43 rows. It adds Claude Opus 5.5 (38.2, rank 1), GPT-6 Astra (34.2), GPT-6 Sol (32.2), Muse Spark 1.3 (32.2) and GPT-6 Luna (25.0), all recorded here. Claude Opus 5's row changed: Snorkel's data notes say the upstream source re-published every Opus 5 figure on 2026-09-24, and this site's 27.0 (Max, score 49.0, est. cost $7,239, recorded 2026-08-20) is gone, replaced by 32.2 at High (score 55.9, est. cost $1,108) and 30.9 at Max; the eight other rows this site had recorded are unchanged. Rows run at several efforts now carry per-effort results, and the board shows each row's best effort by default. This site records that best-effort value, names the effort, and lists the per-effort figures in the row notes, since a best-effort pair can mix tiers (Claude Opus 5's best is High, Opus 5.5's is Max). The board has no row for nineteen tracked models: fourteen with no figure here at all, including Claude Fable 5.1, Claude Sonnet 5.5, Grok 4.7 and GLM-5.3, and five whose figure here is the vendor's own. The board's only DeepSeek V4 Pro row (12.4, OpenClaw) names no build; this site has treated it as a pre-0813 run since August and still excludes it.

Benchmarks related to Agents' Last Exam

Is the current lead on Agents' Last Exam real?

Real gapClaude Opus 5.5 leads by 4.0 points on Agents' Last Exam — outside the noise band, worth acting on.

Written up in full, for models on this Agents' Last Exam board: Gemini 3.1 Pro Preview vs Kimi K3 · Claude Opus 4.8 vs Gemini 3.1 Pro Preview — every benchmark each pair shares, not just Agents' Last Exam.

Agents' Last Exam leaderboard: current scores

Agents' Last Exam’s 20 tracked scores were each verified against their source between 2026-08-17 and 2026-10-01.

Agents' Last Exam scores
ModelAgents' Last Exam score
Claude Opus 5.5
GPT-6 Astra
Claude Opus 5
GPT-6 Sol
Muse Spark 1.3
GPT-5.6 Sol
GPT-5.6 Luna
Kimi K3
Claude Opus 4.8
Qwen3.8-Max
Claude Fable 5
GPT-6 Luna
GLM-5.2
Gemini 3.1 Pro Preview
DeepSeek V4.1 Flash
MiMo-V2.6-Pro
MiMo-V2.6-Flash
GLM-5.3-Flash
DeepSeek V4 Pro (0813)
Qwen3.8-Flash-Next

Agents' Last Exam FAQ

What does Agents' Last Exam measure?

Real, economically valuable professional work — video editing, CAD, manufacturing simulation — done end-to-end. The benchmark has 1000 items.

How big does a gap on Agents' Last Exam have to be to mean anything?

At least 3.2 points. Across 1000 items, a smaller gap sits inside the spread that repeat runs produce on their own, so we call it a tie — provided both scores are independent runs; if either side is a vendor claim, the row is labelled unverified instead.

Can you trust Agents' Last Exam scores?

We grade it A and list it as current — it still tells frontier models apart across 1000 items. A gap of at least 3.2 points clears that noise band, but it only counts as a real gap once both scores come from independent runs.

Source: https://arxiv.org/abs/2606.05405