Methodology
This page explains exactly how every label on this site is decided. If a rule isn’t written here, it isn’t a rule.
The five signal labels
- ✅ Real gap — the difference is bigger than the benchmark’s noise band, and at least one score is independently verified (not just vendor-reported).
- ⚖️ Tie — the difference is smaller than the benchmark’s meaningful gap (see below). Statistically indistinguishable. Ignore the ranking.
- 🔧 Setup-dependent — the two scores come from different setups (with tools vs without, different harnesses). Not a fair comparison; the “winner” can flip depending on setup.
- ⚠️ Unverified — the gap looks real, but both scores are self-reported by the vendor with no independent replication yet.
- 🚱 Tainted — the benchmark itself is saturated or has known contamination issues. We don’t trust rankings on it regardless of the gap size.
What “meaningful gap” means
Every benchmark has a meaningful_gap value: roughly, the smallest score difference that isn’t explained by run-to-run noise given the benchmark’s sample size. We start from a simple statistical floor (a smaller test set needs a bigger gap to mean anything) and adjust down for benchmarks with stronger reliability track records. It is not a precise confidence interval — it’s a conservative, published threshold you can check us on.
Benchmark trust grades
We grade each benchmark A through D on how much weight its scores deserve: contamination risk, sample size, and how much of the field self-reports versus gets independently verified.
| Benchmark | Grade | Items | Status |
|---|---|---|---|
| 🧠 Reasoning | B | 2500 | current |
| 💻 Terminal ops | B | 89 | current |
| 🛠️ Long-horizon coding | B | 113 | current |
| 🧰 Multi-tool chores | B | 108 | current |
| 💼 Professional work | A | 1000 | current |
| 🐛 Bug fixing | D | 500 | saturated |
| 🔬 Expert science Q&A | C | 198 | aging |
| ⌨️ Contest coding | B | 1055 | current |
A = designed to resist contamination (e.g. a living benchmark refreshed on a schedule). B = reliable, mostly vendor-reported, no major known issues. C = usable but small sample size or nearing saturation — treat close results with extra skepticism. D = saturated or known contamination — rankings here are treated as unreliable regardless of gap size.
Why we never publish a single composite score
Combining several benchmarks into one weighted number requires picking which version of a multi-setting benchmark counts — and that choice alone can flip the ranking. On Humanity’s Last Exam, DeepSeek V4 Pro trails Claude Opus 4.8 without tools and leads with tools. A composite score would have to silently pick one of those numbers, dress a scaffolding choice up as model quality, and call it precision it doesn’t have.
Instead, we publish tiers and per-benchmark verdicts, not rankings with false precision. A tier board — grouping models that are statistically indistinguishable on trust-graded benchmarks, with zero weighting — is in progress; this page will document its exact rule before it goes live.
Data sourcing
Every score on this site carries a source link and a self-reported/independent tag. If we can’t find a real, checkable source, the field says “Not sourced yet” — we don’t estimate or fill gaps with a plausible-looking number.