AI benchmarks explained
Model leaderboards rank the players. This table grades the referees: each of the 12 benchmarks we track gets a trust grade and a noise threshold, and every comparison on this site inherits both. Most are grouped by what they actually ask a model to do — reasoning, coding, agentic work — because a score only transfers to the kind of task it was measured on. A fourth group, composite, holds the rare benchmark that averages several of those at once and so doesn’t belong to any single one.
How to read this table
Trust grade (A–D) answers one question: when this benchmark says two models differ, how seriously should you take it? An A means the benchmark actively defends itself — refreshed items, contamination controls, a sample large enough that small gaps still carry signal. B is the workhorse grade: no known structural problem, but nothing an A has to defend it either. C means age or a thin sample has started eroding what a score can tell you. D means the ceiling has been reached or the test set has leaked — scores still get printed everywhere, but once a benchmark’s status reads saturated or contaminated, every comparison on it here is labeled Tainted rather than ranked.
Meaningful gap is the point difference below which two scores are statistically indistinguishable on that benchmark — mostly a function of sample size. An 89-item suite simply cannot resolve a 3-point difference the way a 2,500-item one can, which is why the thresholds in this table range from ±2 to over ±10. Between two independently-run scores, any gap under the threshold gets called a tie here, whatever order the leaderboards print.
Category groups benchmarks by what they ask a model to do, because that is what determines whether a score transfers to your work. Reasoning tests hand the model a question with a known answer; coding tests grade a patch by running the project’s own tests; agentic tests turn it loose on a long task with real tools and score the end state. A model can be genuinely excellent at one and ordinary at another — comparing across categories is the single most common way benchmark tables mislead.
The grouping is by what form the work takes, not by which subjects it touches — Terminal-Bench’s tasks span software engineering, sysadmin, data processing and security, but every one of them is graded the same way: turn a model loose on a real terminal and check the end state. That is what makes it agentic, regardless of how many fields its task prompts happen to mention.
Status tracks where the benchmark is in its life: current ones still separate frontier models, aging ones are losing resolution as the field climbs toward their ceiling, and saturated or contaminated ones no longer produce comparisons worth trusting at all.
The full grading criteria live in the methodology; the grades and thresholds themselves ship as CC BY JSON on the open data page. Each benchmark name above links to a deeper page: what it measures, its caveats, and current scores — or put any of it to work in the comparison tool.
| Benchmark | Trust grade | Meaningful gap | Items | Status |
|---|---|---|---|---|
| ReasoningClosed-ended questions with a known answer — what the model knows and can work out on its own. | ||||
| Humanity's Last Exam Reasoning Expert-written questions across dozens of subjects, built to be hard enough that today's best models still fail most of them. | B | ±2 | 2500 | current |
| GPQA Diamond Expert science Q&A PhD-level science questions written to be 'Google-proof' — hard even for skilled non-experts with internet access. | D | ±7.2 | 198 | saturated |
| ARC-AGI-2 Compositional visual reasoning Fluid, compositional visual reasoning: inferring a hidden transformation rule from a handful of example input/output grids and applying it exactly to a new one. | B | ±9.2 | 120 | current |
| HMMT Feb 2026 Competition mathematics Whether a model can solve HMMT Feb 2026's short-answer competition math problems — pre-college-level algebra, number theory, combinatorics, and geometry, each with one objectively-checkable numeric or symbolic answer. | C | ±17.5 | 33 | contaminated |
| CodingWrite or repair real code, graded by whether the tests pass. | ||||
| DeepSWE Long-horizon coding Original, never-before-public software engineering tasks, built so agents can't just recall memorized GitHub fixes. | B | ±9.5 | 113 | current |
| SWE-bench Verified Bug fixing Gives an agent a real closed GitHub issue and checks whether its patch fixes it and passes the project's real test suite. | D | ±4.5 | 500 | saturated |
| LiveCodeBench Contest coding Coding contest problems dated after each model's training cutoff, so scores can't be inflated by memorization. | D | ±3.1 | 1055 | saturated |
| AgenticLong multi-step work with real tools, graded on what the run produces — an end state or a final figure — not on a single completion. | ||||
| Terminal-Bench 2.1 Terminal ops Whether an AI agent can actually operate a real command line to debug code, administer systems, and fix security issues. | B | ±10.6 | 89 | current |
| Toolathlon-Verified Multi-tool chores Realistic multi-step chores that require an agent to juggle many real software tools over a long session, not just answer questions. | B | ±9.7 | 108 | current |
| Agents' Last Exam Professional work Real, economically valuable professional work — video editing, CAD, manufacturing simulation — done end-to-end. | A | ±3.2 | 1000 | current |
| AA-AnalystAgent Spreadsheet & document analysis Whether an agent can answer the kind of quantitative question a business or data analyst faces day to day — working from a folder of real spreadsheets and documents, running Python in a sandbox, and returning the one right value on all five of five attempts. | B | ±11.2 | 80 | current |
| CompositeOne score averaged across several different kinds of task — no single execution form to point to. | ||||
| LiveBench Composite score across 7 domains LiveBench's tracked "Overall" score is a composite average across seven independently-graded categories — reasoning, coding, agentic coding, math, data analysis, language, and instruction following — each scored against an objective, verifiable ground-truth answer rather than any single skill. | A | ±2.7 | 1436 | current |
Further reading on AI benchmark trust
- GPQA Diamond leaderboard 2026 — We graded our own GPQA Diamond leaderboard a D and stopped ranking it. Here is the data behind that call, and what changed across all 253 model pairs.