Open data

Every number on this site is a view of these five files. Licensed CC BY 4.0 — use it, cite us.

34 models · 12 benchmarks · 260 sourced scores (217 independent, 43 vendor-reported — 83% independent) · 5 head-to-head verdicts · 45 release events · latest observation 2026-10-01

  • models.json — Model specs and official list pricing
  • benchmarks.json — Benchmark metadata, trust grades, meaningful-gap thresholds
  • scores.json — Every sourced score, with source URL and self-reported/independent tag
  • pairs.json — Curated head-to-head verdicts
  • releases.json — Model release / price-change / deprecation events with a one-line verdict each

What makes this dataset different

Provenance is a first-class column, not a footnote. Every score row records where the number was read (source_url) and who produced it (source_type: the vendor’s own launch chart, or an independent re-run). That distinction drives every label on the site: on a healthy benchmark, a gap built on a vendor claim is marked Unverified no matter how large it is — it can never be called a Real gap. It’s the field that decides whether a chart-topping number has ever been reproduced by anyone.

The files above are not an export — they are the same source-of-truth JSON the site itself is built from, mirrored verbatim at build time. When a score changes, the pages and the download change together; there is no second copy to drift.

Field guide

scores.json — 260 rows

model_id / benchmark_idForeign keys into models.json and benchmarks.json
variantEvaluation setup: no_tools, with_tools, or default. Scores from different variants are never compared — the site labels such pairings Setup-dependent
valueThe score as published by the source, unmodified
source_urlThe page the number was read from
source_typeindependent = someone other than the model’s vendor ran the eval; self-reported = the vendor’s own launch number. A self-reported side can never earn a comparison the Real gap or Tie label
date_observedWhen we recorded the value (YYYY-MM-DD)
notesCaveats worth knowing before citing the row — harness details, superseded values, ambiguities

benchmarks.json — 12 rows

measuresWhat the benchmark actually tests, in one plain sentence
n_itemsNumber of items in the test set — the denominator behind the noise math
meaningful_gapPoints of difference below which a gap is treated as noise on this benchmark
trust_gradeA–D: how much a score on this benchmark should move your beliefs (see /methodology)
lifecyclecurrent, aging, saturated, or contaminated. Saturated and contaminated benchmarks mark every comparison Tainted
variantsWhich evaluation setups exist for this benchmark (e.g. no_tools / with_tools)

models.json — 34 rows

lab / release_date / licenseWho shipped it, when, and proprietary vs open-weights
price_in / price_outOfficial list price in USD per 1M tokens (input / output); null where no per-token list price exists. price_source_url links the vendor’s pricing page
context_windowMaximum context in tokens; null where unpublished
lifecycle / superseded_bycurrent, superseded, or deprecated — superseded models stay in the data so old comparisons remain reproducible

pairs.json carries the editorial head-to-head verdicts, including a verdict_basis hash of the signal table each verdict was written against — the build fails if the underlying scores drift under the prose. releases.json is the event log behind the new-models tracker: date, event type (release / price-change / deprecation), source URL, and a one-line note on whether it matters.

Access

Fetch the files with curl, a cron job, or an agent — no API key, no login, and robots.txt allows every user-agent.

curl https://themodelgap.com/data/scores.json

Two real rows from that file, unmodified:

[
  {
    "model_id": "deepseek-v4-pro",
    "benchmark_id": "hle",
    "variant": "no_tools",
    "value": 41,
    "source_url": "https://artificialanalysis.ai/evaluations/humanitys-last-exam",
    "source_type": "independent",
    "harness": "artificial-analysis",
    "date_observed": "2026-08-17",
    "notes": "AA's own run of 'DeepSeek V4 Pro 0813 (Reasoning, Max Effort)' — text-only, no tools. Replaces the launch chart's self-reported 42.7. AA's run of the older 0424 V4 Pro scored 37.5; don't conflate."
  },
  {
    "model_id": "deepseek-v4-pro",
    "benchmark_id": "hle",
    "variant": "with_tools",
    "value": 60,
    "source_url": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813",
    "source_type": "self-reported",
    "harness": "vendor-self-report",
    "date_observed": "2026-08-13",
    "notes": ""
  }
]

There’s also a machine-readable site guide at /llms.txt and an RSS feed at /feed.xml. If you build something on this data, tell us — we’re happy to link back.

How to cite

Plain text: Data: The Model Gap (themodelgap.com), CC BY 4.0. For a specific number, prefer citing the row’s own source_url alongside us — the original evaluator deserves the credit for running it.

@misc{themodelgap,
  author = {{The Model Gap}},
  title  = {AI model benchmark scores with provenance},
  year   = {2026},
  url    = {https://themodelgap.com/data},
  note   = {CC BY 4.0; each score carries a source URL
            and a self-reported/independent tag}
}

How the labels and trust grades are computed from these fields: methodology.