Methodology

This page explains exactly how every label on this site is decided. If a rule isn’t written here, it isn’t a rule.

The five signal labels

They are checked in this order, and the first one that applies is the one you see. A gap has to clear every earlier test before its size is even looked at. Setup-dependent appears twice in the flow: variant differences (tools, prompts) are caught before the provenance check, harness differences after it — either way the label is the same.

  • Tainted — the benchmark itself is saturated or has known contamination issues. We don’t trust rankings on it regardless of the gap size.
  • Setup-dependent — the two scores come from different setups: a different variant (with tools vs without), caught before the provenance check, or a different evaluation harness, caught after it. Not a fair comparison; the “winner” can flip depending on setup.
  • Unverified — either score is the vendor’s own claim rather than an independent run. One vendor-reported side is enough: a lab has every incentive to publish a flattering number and none to publish an unflattering one, so the risk runs one way. That applies to the close calls too — a tie between a measurement and a claim isn’t a tie yet.
  • Tie — both scores are independent runs, and the difference is smaller than the benchmark’s meaningful gap (see below). Statistically indistinguishable. Ignore the ranking.
  • Real gap — both scores are independent runs and the difference clears the benchmark’s noise band. This is the only label that says one model is actually ahead.

What “meaningful gap” means

Every benchmark has a meaningful_gap value: roughly, the smallest score difference that isn’t explained by run-to-run noise given the benchmark’s sample size. It comes from one formula applied identically to all 12: 100 ÷ √n, rounded up to one decimal — two standard errors on a pass/fail test set of n items. A 2,500-item benchmark earns a 2.0-point band; an 89-item one needs 10.6. It is not a precise confidence interval, and no benchmark gets a discount for having a good reputation — it is one published number per benchmark that you can recompute yourself from the item count in the table below.

Benchmark trust grades

We grade each benchmark A through D on how much weight its scores deserve: contamination risk, sample size, and how much of the field self-reports versus gets independently verified.

BenchmarkCategoryGradeItemsStatus
HLEReasoningB2500current
Terminal-Bench 2.1AgenticB89current
DeepSWECodingB113current
Toolathlon-VerifiedAgenticB108current
Agents' Last ExamAgenticA1000current
SWE-bench VerifiedCodingD500saturated
GPQA DiamondReasoningD198saturated
LiveCodeBenchCodingD1055saturated
ARC-AGI-2ReasoningB120current
LiveBenchCompositeA1436current
HMMTReasoningC33contaminated
AnalystAgentAgenticB80current

A = designed to resist contamination (e.g. a living benchmark refreshed on a schedule). B = reliable, mostly vendor-reported, no major known issues. C = usable but small sample size or nearing saturation — treat close results with extra skepticism. D = saturated or known contamination — rankings here are treated as unreliable regardless of gap size.

Why we never publish a single composite score

Combining several benchmarks into one weighted number requires picking which version of a multi-setting benchmark counts — and that choice alone can flip the ranking. On Humanity’s Last Exam, DeepSeek V4 Pro trails Claude Opus 4.8 without tools and leads with tools. A composite score would have to silently pick one of those numbers, dress a scaffolding choice up as model quality, and call it precision it doesn’t have.

Instead, we publish tiers and per-benchmark verdicts, not rankings with false precision. A tier board — grouping models that are statistically indistinguishable on trust-graded benchmarks, with zero weighting — is in progress; this page will document its exact rule before it goes live.

Data sourcing

Every score on this site carries a source link and a self-reported/independent tag. If we can’t find a real, checkable source, the field says “Not sourced yet” — we don’t estimate or fill gaps with a plausible-looking number.

The rules on this page are the how. For the why — what this site is for, who runs it, how corrections are handled and how it is funded — read about The Model Gap. Every number here is open: download the full dataset and check any of it yourself.