About

Every week a new AI model launches with a table of benchmark scores, and every leaderboard ranks them 1, 2, 3.

Most of those rankings are noise.

A 0.5-point gap on a benchmark with a few hundred questions is a coin flip. A model can lose a benchmark without tools and win it with them — same model, opposite verdict. Scores get reported by the labs themselves, on harnesses nobody else can reproduce.

The Model Gap reads each release and tells you one thing: which gaps are real, and which are just noise.

No composite score. No pretending 0.5 points means something.

The labels

How we work

Every number on this site carries a source link and a self-reported or independent tag — nothing gets published without both. Scores come from official vendor material, or from third parties that run the benchmarks themselves (vals.ai, Artificial Analysis, and similar). We don’t run our own evals — every judgment here is about reading existing scores honestly, not producing new ones. See the open data for the full sourced table.

Corrections

We publish corrections in place, dated, rather than quietly editing a number and moving on. Three examples so far:

  • 2026-08-19 — DeepSeek V4 Pro vs. Claude Opus 4.8: an Agents’ Last Exam score was transcribed from the wrong column (25.2 instead of 25.7) — the note explaining the fix is still in the article.
  • 2026-08-19 — DeepSeek V4 Pro and DeepSeek V4 Flash: the two models’ listed prices were on different discount bases (one peak, one off-peak), making Pro look 1.5x Flash’s price instead of the true 3x — both are now shown at the same, undiscounted rate.
  • 2026-08-17 — Self-reported vs. independent AI benchmarks: published with a 69%-vendor-reported figure, then updated the same day after we replaced most of those rows with independent runs — the “Update, later the same day” section at the bottom is the original correction, left in place rather than merged away.

If you find a number that’s wrong, tell us: info@themodelgap.com.

Funding

The Model Gap is unmonetized today — no ads, no sponsored placements, no affiliate links. None of the labs whose models we grade pay us anything. If that changes, one rule is fixed in advance: any inference-related revenue (for example, a future API gateway) would apply the same margin to every model, would never charge for bring-your-own-key usage, and would be disclosed on every page it touches. Pricing a comparison differently depending on which model earns us more would make every verdict on this site untrustworthy, so it’s not a choice we’re leaving open.

Who’s behind this

The Model Gap is a one-person project. There’s no lab-sponsored leaderboard, no SaaS upsell riding on a favorable ranking — which is also why we can say a benchmark result is a coin flip without it costing us anything.

Get each verdict by email: themodelgap.substack.com · Data: open, CC BY · Contact: info@themodelgap.com