GPQA Diamond is saturated — here is what that costs our own table

August 24, 2026 · Updated October 6, 2026

GPQA Diamond is saturated: the 23 frontier models this site tracks independently span only 10.45 points — 96.06% at the top, 85.61% at the bottom — and 22 of the 23 sit inside the benchmark's 7.2-point noise band of the leader, which means no ordering a leaderboard prints on it separates real capability differences from run-to-run variance. So on August 20, 2026, we downgraded our own GPQA Diamond leaderboard to trust grade D, marked it saturated, and turned off ranking for every pair of models that share the benchmark. Not just the close calls. Every pair, all 253 of them.

That is not a hedge. It is a rule enforced in code, and this is the receipt.

The data behind the D

We track 23 models on GPQA Diamond, all independently scored — no vendor self-reports made the cut here (DeepSeek V4.1 Flash, added 2026-09-10, StepFun's Step 5 Preview, added 2026-09-22, and Tencent Hy4 preview, added 2026-09-27, each carry a GPQA Diamond number, but all three are the vendor's own claims, not independent runs, so they sit outside this 23-model set — they would be just as tainted as everyone else here if they were counted, since the saturation rule applies before provenance is even checked, but the pairs-and-spread figures below are scoped to the models this piece can independently verify). Sorted high to low, the spread runs from GPT-6 Astra at 96.06% down to GLM-5.2 at 85.61%: a 10.45-point range across the whole field, with a population standard deviation of about 2.27 points.

Compare that to the benchmark's own noise band. GPQA Diamond has 198 items, and two standard errors on a set that size works out to 7.2 points — the meaningful gap below which a difference is statistically indistinguishable from run-to-run variance. Twenty-two of our 23 tracked models sit within that 7.2-point band of the leader. The one exception, GLM-5.2, isn't isolated either: its closest tracked neighbor by score, Claude Sonnet 5, sits just 3.28 points above it, safely inside the band. The model that comes nearest to actually crossing the threshold against GLM-5.2 without doing so — MiniMax M3 — lands 7.07 points ahead, just 0.13 points short. Every model we track has at least one other tracked model within noise distance of it — a close top, and no clear bottom.

band edge: leader − 7.284%86%88%90%92%94%96%98%GPT-6 Astra 96.06%GLM-5.2 85.61%
The 23 independently scored GPQA Diamond results this piece counts. Shaded region: within 7.2 points of the leader — the band inside which a difference is statistically indistinguishable. GLM-5.2 (accent dot) is the only tracked model outside it.

GPQA Diamond leaderboard 2026: the 23 independently scored models

The highest independent GPQA Diamond score tracked here is 96.06%, and the whole field sits within 10.45 points of it. Below is every model this piece counts, sorted by score for reading. The last column is what our compareScores() rule would have returned for each model against GLM-5.2 — the bottom of the table and the only model outside the leader's 7.2-point band — before the August 20 downgrade. Today every one of these 253 pairs reads Tainted, so the order here is a fact about the numbers, not a verdict about the models.

ModelGPQA DiamondEvaluatorObservedRead against GLM-5.2, before the downgrade
GPT-6 Astra96.06%Artificial Analysis2026-09-04Setup-dependent (different harness)
Gemini 3.1 Pro Preview95.45%vals.ai2026-08-17Real gap (+9.84)
GPT-5.6 Sol95.2%vals.ai2026-08-17Real gap (+9.59)
Grok 4.694.7%vals.ai2026-08-17Real gap (+9.09)
Gemini 3.8 Flash94.44%vals.ai2026-09-03Real gap (+8.83)
Gemini 3.7 Flash · previous version93.94%vals.ai2026-08-17Real gap (+8.33)
Qwen3.8-Max93.69%vals.ai2026-08-17Real gap (+8.08)
Muse Spark 1.393.5%Artificial Analysis2026-09-29Setup-dependent (different harness)
Claude Opus 5 · previous version93.43%vals.ai2026-08-20Real gap (+7.82)
Claude Fable 5.193.43%vals.ai2026-09-02Real gap (+7.82)
Claude Fable 5 · previous version93.18%vals.ai2026-08-17Real gap (+7.57)
Kimi K392.93%vals.ai2026-08-17Real gap (+7.32)
MiniMax M392.68%vals.ai2026-08-26Tie (+7.07, inside the 7.2-point band)
DeepSeek V4 Pro (0813)92.42%vals.ai2026-08-17Tie (+6.81, inside the 7.2-point band)
Claude Opus 4.8 · previous version92.42%vals.ai2026-08-17Tie (+6.81, inside the 7.2-point band)
Qwen3.8-Flash-Next92.3%Artificial Analysis2026-08-28Setup-dependent (different harness)
GLM-5.391.7%Artificial Analysis2026-08-20Setup-dependent (different harness)
GPT-5.6 Luna91.67%vals.ai2026-08-26Tie (+6.06, inside the 7.2-point band)
GLM-5.3-Flash91.2%Artificial Analysis2026-08-26Setup-dependent (different harness)
Muse Spark 1.2 · previous version90.4%Artificial Analysis2026-08-26Setup-dependent (different harness)
DeepSeek V4 Flash (0731) · previous version89.9%vals.ai2026-08-17Tie (+4.29, inside the 7.2-point band)
Claude Sonnet 5 · previous version88.89%vals.ai2026-08-17Tie (+3.28, inside the 7.2-point band)
GLM-5.2 · previous version85.61%vals.ai2026-08-17—

Eight of the 23 are previous versions, kept for comparison. The six Setup-dependent rows are the models whose tracked score comes from Artificial Analysis rather than vals.ai — a different harness from GLM-5.2's, which is why they never entered the real-gap count even when their raw score, like GPT-6 Astra's, topped the table.

What "saturated" actually costs

Before the downgrade, our own signal logic — the same compareScores() function that grades every pairwise matchup on the site — would have called 10 of these 253 pairs a confirmed Real gap, 141 a Tie, and 102 Setup-dependent (every one of those 102 involves GLM-5.3, GLM-5.3-Flash, Muse Spark 1.2, Muse Spark 1.3, Qwen3.8-Flash-Next, or GPT-6 Astra — the six models whose tracked GPQA score comes from a different evaluation harness than the other 17). Zero would have come back Unverified, because every score here is independently sourced.

Those 10 "real" gaps weren't spread across the field — they were one shape repeated ten times: every model that shares GLM-5.2's own vals.ai evaluation harness and beats it by more than 7.2 points — not simply "the top ten by score." As of this refresh the outright #1 model on the whole board, GPT-6 Astra (96.06%, Artificial Analysis), sits outside this group entirely on harness grounds — the same exclusion that already applied to Muse Spark 1.3 (93.5%), just now happening to the single highest score on the table — while Kimi K3 (92.93%) sits inside the group despite ranking twelfth overall. The group includes Claude Fable 5.1 (tied with Claude Opus 5 at 93.43%) and Gemini 3.8 Flash (94.44%, slotting in between Grok 4.6 and Gemini 3.7 Flash). (One of those ten is Claude Fable 5's 93.18% — a number that carries its own caveat on our benchmark page: it counts refusal-triggered fallback answers as correct, and scoring refusals as failures instead drops it to 55.56%. Even the numbers feeding our own "real" column aren't uniformly clean.) Three others land just short of the threshold and so would have read as ties under the old grading: MiniMax M3 clears GLM-5.2 by 7.07 points, DeepSeek V4 Pro and Claude Opus 4.8 by 6.81 each — wide absolute margins that the 7.2-point band still swallows.

After the downgrade, all 253 pairs read Tainted. Not because we recalculated and found the gaps smaller — the lifecycle check in our code runs before the gap is even computed. For a saturated benchmark, compareScores() returns Tainted on its first line; the point difference between two models never gets touched. Ten pairs that used to clear our own bar for "real" now carry the identical label as a pair separated by 0.00 points.

Why a ceiling does this to a benchmark

Once most of a field clusters within a couple of points of the maximum score, a benchmark stops discriminating between models and starts measuring who got lucky on the hardest handful of questions. Epoch AI's "GPQA Diamond: What's left?" put numbers on that risk: it flagged the weakest-performing 40 of the 198 questions for review, went deep on the six most extreme cases — items where every model tested scored under 5% — and judged more than a third of that handful flawed, which it scales up to a benchmark-wide error estimate near 8%. Epoch's own conclusion is more measured than "the benchmark is broken": it still counts at least 90% of GPQA Diamond as sound. But a one-in-twelve error rate sitting on top of a field already bunched within single digits of a 100% ceiling is exactly the combination that turns a ranking into a coin flip.

What the outside world shows, as of today

Checking the live leaderboard rather than trusting our own capture: re-checked August 27, 2026, vals.ai's GPQA Diamond tracker lists 135 models, 24 of them scoring 90% or higher — unchanged from this piece's original August 24 check, and the same 24-model count we cited from our August 17 snapshot of 133 models. The leader hasn't moved through any of it: Gemini 3.1 Pro Preview, still 95.45%. vals.ai's own read on that density is blunt — a high score, in their words, "no longer meaningfully distinguishes frontier models."

Contamination evidence isn't static either, though it isn't a clean upward line. The n-gram audit we cite on our benchmark page (Xu et al., EMNLP 2025) checked GPQA's original 448-question set against six Common Crawl snapshots and found the contaminated-item count bouncing between 3 and 12 across them — not a steady climb, but net higher at the end than at the start: 4 flagged items in the earliest snapshot tested (CC-2025-05), dipping as low as 3 partway through, then 12 in the most recent one (CC-2025-26). A benchmark whose contamination count can double or triple between snapshots, in either direction, is its own argument for a D — separate from the ceiling problem.

One thing worth being precise about: this isn't just a polite ask. GPQA's dataset is gated on Hugging Face — downloading it requires clicking through an agreement stating "you agree to NOT reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model training corpora." That's a real condition of access, even though it isn't a copyright license in the legal sense. We're not quoting a test item here either way.

The mechanism, not a one-off call

None of this is a manual override applied to one page. compareScores() checks a benchmark's lifecycle before it looks at anything else — before matching variants, before checking whether either score is vendor-reported, before comparing harnesses, before any gap is measured against the 7.2-point threshold. A saturated or contaminated benchmark exits the function on that first check, every time, for every pair. The full decision order, and the other four labels a comparison can land on, is laid out on our methodology page.

So what do you read instead

Not "GLM-5.2 is bad at science questions" — an 85.61% on a Google-proof PhD-level exam is a real result, and the distance to the field's leader sits inside what repeat runs alone would produce. If you need to compare two models' scientific reasoning today, GPQA Diamond isn't where that comparison should happen; benchmarks with a higher trust grade still separate models cleanly. If you want the underlying numbers yourself, sources and dates included, they're open in our data export.

GPQA Diamond: quick questions

Is GPQA Diamond saturated?

Yes. Twenty-two of the 23 independently-scored models tracked here sit inside the benchmark's 7.2-point noise band of the leader, and the whole field spans just 10.45 points — 96.06% at the top (GPT-6 Astra) down to 85.61% (GLM-5.2). This site graded the benchmark D and turned ranking off for every pair on 2026-08-20.

What does saturated mean for a benchmark?

The instrument's range is used up: tracked models cluster so tightly near the ceiling that a leaderboard's ordering no longer separates capability from run-to-run variance. For GPQA Diamond the meaningful-gap threshold implied by its 198 items is ±7.2 points, and almost every tracked pair falls inside it.

How many questions are in GPQA Diamond?

198 in the Diamond subset. That item count is where the 7.2-point noise band comes from — two standard errors on a set of that size — and it is why a three-point difference between two models on this board says nothing.

Who has the highest GPQA Diamond score?

GPT-6 Astra at 96.06%, an Artificial Analysis run recorded 2026-09-04 — the outright leader by raw score, and per this piece's revision log, the first number-one that sits in the setup-dependent minority by harness rather than the real-gap cluster.

Which model has the lowest tracked GPQA Diamond score?

GLM-5.2 at 85.61% — the only model outside the leader's 7.2-point band. Even it is not isolated: its nearest tracked neighbor, Claude Sonnet 5, sits 3.28 points above, safely inside the band.

Why are all 253 pairs labeled tainted?

Because the saturation rule fires before any gap arithmetic: once a benchmark is graded saturated here, every pair sharing it is labeled Tainted however large the difference looks — not just the close calls. That is the rule enforced in code on 2026-08-20.

Will new models be added to GPQA Diamond leaderboards?

Not by the two main independent sources. vals.ai has archived its board — "we no longer run this benchmark on new model releases" — and Artificial Analysis dropped GPQA Diamond from its Intelligence Index in v4.2 (September 2026), publishing no figure for any model released after 2026-09-11. Vendor self-reported numbers still appear; they are excluded from this table.

Did any model come close to crossing the noise band against GLM-5.2?

MiniMax M3 came nearest before the downgrade froze the calls: 7.07 points ahead of GLM-5.2, just 0.13 short of the 7.2-point band — the largest gap any tracked model posted while staying inside the band.

Revision log

  • Updated August 27, August 28, September 2, twice more on September 3, and again on September 4, 2026: this piece's counts have now been recomputed six times as our GPQA Diamond table grew from 14 to 18, then 19, then 20, then 21, then 22, then 23 models (91 → 153 → 171 → 190 → 210 → 231 → 253 pairs). The 21st was Gemini 3.8 Flash at 94.44% (Setup-dependent harness aside, it joined the real-gap cluster, moving that count 9 → 10); the 22nd, added the same day, was Muse Spark 1.3 at 93.8% (Artificial Analysis) — its harness puts it in the Setup-dependent minority against most of the field, so that refresh moved the pair and Setup-dependent counts but left the real-gap count at 10. The 23rd, GPT-6 Astra at 96.06% (also Artificial Analysis), is now the outright leader by score — and the first time the raw #1 model on the whole board sits in the Setup-dependent minority, not the real-gap cluster: the real-gap count holds at 10 a second time. The shape underneath still has not changed once: still one pattern repeated, same 10 models, now a 10.45-point spread. The pair counts, standard deviation, setup-dependent tally and nearest-miss model in the body above are current; the August 20 decision and its reasoning are unchanged.
  • Updated again September 10, 2026: DeepSeek V4.1 Flash shipped with a self-reported (not independent) GPQA Diamond score — it does not join the 23-model tracked set in the body above, since this piece only counts independently-scored models, but a scoping note was added to say so explicitly rather than leave the 253-pair figure looking exhaustive of every model with any GPQA number at all. No pair, spread, or standard-deviation figure changed.
  • Updated again September 22, 2026: StepFun's Step 5 Preview arrived with a second self-reported GPQA Diamond number (93.5); the scoping note now names both, and again no pair, spread, or standard-deviation figure changed.
  • Updated again September 27, 2026: Tencent Hy4 preview added with a third self-reported GPQA Diamond number (92.3); same treatment, no figure changed.
  • Updated again September 29, 2026: the full table of the 23 independently scored models this piece counts is now on the page, each with its pre-downgrade read against GLM-5.2 — previously the piece quoted individual scores in prose but carried no complete table. Same day: Artificial Analysis's Muse Spark 1.3 (max) row moved from 93.8 to 93.5 on a re-read of its board, so the table and the standard deviation (2.28 → 2.27) follow it; no pair count changed. Also that day, a source note: vals.ai — behind 17 of the 23 rows here — has archived its GPQA Diamond board ("Since performance on this benchmark has saturated, we no longer run this benchmark on new model releases"), and Artificial Analysis dropped GPQA Diamond from its Intelligence Index in v4.2 (September 2026) and has published no GPQA figure for any model released after 2026-09-11. As things stand, no new model will join this table from either source.

Sources: 23 tracked GPQA Diamond scores and their provenance from our own scores data, cross-checked against the vals.ai GPQA Diamond leaderboard (checked August 24, re-checked August 27, 2026). Question-quality analysis from Epoch AI, "GPQA Diamond: What's left?" (May 30, 2025). Contamination figures from Xu et al., "Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-Index" (EMNLP 2025, Table 7), verified against the paper's full text. Dataset access terms from the GPQA dataset card on Hugging Face. The rule behind every label in bold above is written out in the methodology.