The Model Gap

Methodology

This page explains exactly how every label on this site is decided. If a rule isn’t written here, it isn’t a rule.

The five signal labels

  • ✅ Real gap — the difference is bigger than the benchmark’s noise band, and at least one score is independently verified (not just vendor-reported).
  • ⚖️ Tie — the difference is smaller than the benchmark’s meaningful gap (see below). Statistically indistinguishable. Ignore the ranking.
  • 🔧 Setup-dependent — the two scores come from different setups (with tools vs without, different harnesses). Not a fair comparison; the “winner” can flip depending on setup.
  • ⚠️ Unverified — the gap looks real, but both scores are self-reported by the vendor with no independent replication yet.
  • 🚱 Tainted — the benchmark itself is saturated or has known contamination issues. We don’t trust rankings on it regardless of the gap size.

What “meaningful gap” means

Every benchmark has a meaningful_gap value: roughly, the smallest score difference that isn’t explained by run-to-run noise given the benchmark’s sample size. We start from a simple statistical floor (a smaller test set needs a bigger gap to mean anything) and adjust down for benchmarks with stronger reliability track records. It is not a precise confidence interval — it’s a conservative, published threshold you can check us on.

Benchmark trust grades

We grade each benchmark A through D on how much weight its scores deserve: contamination risk, sample size, and how much of the field self-reports versus gets independently verified.

BenchmarkGradeItemsStatus
🧠 ReasoningB2500current
💻 Terminal opsB89current
🛠️ Long-horizon codingB113current
🧰 Multi-tool choresB108current
💼 Professional workA1000current
🐛 Bug fixingD500saturated
🔬 Expert science Q&AC198aging
⌨️ Contest codingB1055current

A = designed to resist contamination (e.g. a living benchmark refreshed on a schedule). B = reliable, mostly vendor-reported, no major known issues. C = usable but small sample size or nearing saturation — treat close results with extra skepticism. D = saturated or known contamination — rankings here are treated as unreliable regardless of gap size.

Why we never publish a single composite score

Combining several benchmarks into one weighted number requires picking which version of a multi-setting benchmark counts — and that choice alone can flip the ranking. On Humanity’s Last Exam, DeepSeek V4 Pro trails Claude Opus 4.8 without tools and leads with tools. A composite score would have to silently pick one of those numbers, dress a scaffolding choice up as model quality, and call it precision it doesn’t have.

Instead, we publish tiers and per-benchmark verdicts, not rankings with false precision. A tier board — grouping models that are statistically indistinguishable on trust-graded benchmarks, with zero weighting — is in progress; this page will document its exact rule before it goes live.

Data sourcing

Every score on this site carries a source link and a self-reported/independent tag. If we can’t find a real, checkable source, the field says “Not sourced yet” — we don’t estimate or fill gaps with a plausible-looking number.