GPQA Diamond
PhD-level science questions written to be 'Google-proof' — hard even for skilled non-experts with internet access.
What does a GPQA Diamond task look like?
One item is a single graduate/PhD-level multiple-choice question in biology, physics, or chemistry, written by a subject-matter expert, with four answer options and one unambiguous correct choice. Items are deliberately constructed to survive an open-internet search: in the benchmark's own validation study, skilled non-expert annotators given the question plus 30+ minutes of unrestricted web access answered only about 34% correctly, versus about 65% for on-domain PhD-level experts. GPQA's publishers ask that the exact wording of individual questions not be reproduced in plaintext or images online (to limit training-data leakage), so no example item is quoted here — only the format is described.
How many points on GPQA Diamond is real?
On GPQA Diamond the 7.2-point noise band is beside the point: the benchmark is saturated, so every comparison across its 198 items is labeled tainted instead of being measured against a threshold.
Grade D: GPQA Diamond is saturated or has known contamination issues — its rankings are unreliable at any gap size, 198 items or not.
23 independent GPQA Diamond scores, no noise band drawn: the benchmark is saturated, so every GPQA Diamond head-to-head is labelled tainted rather than measured against its 7.2-point threshold. How tightly the GPQA Diamonddots bunch above is the reason for that call. 3 vendor-reported GPQA Diamond scores are not plotted — see the table below.
Can you trust GPQA Diamond? Caveats
Most frontier models score near the maximum on GPQA Diamond — this benchmark no longer separates them well.
198 questions, the highest-quality cut of an original 448-question set built by GPQA's authors — the subset where both expert validators answered correctly and at most one of three skilled non-experts with internet access did, which selects for reliably Google-proof items rather than simply the hardest ones (items the experts themselves missed are excluded by that rule) (David Rein and coauthors, including NYU's Julian Michael and Samuel Bowman; "GPQA: A Graduate-Level Google-Proof Q&A Benchmark," arXiv:2311.12022, submitted Nov 2023). Top models now cluster within 1-3 points of each other near the ceiling — differences here are increasingly noise, not signal.
Scoring mechanics
Plain accuracy on 4-option multiple-choice questions (25% random-guess floor). The model's selected option is matched against a single fixed correct answer — there is no LLM judge grading free-form responses. Independent evaluators such as vals.ai report both zero-shot and few-shot chain-of-thought accuracy.
Saturation evidence
On vals.ai's independent leaderboard (the same ranking basis used for the scores tracked on this page — 133 models as of this site's 2026-08-17 capture date, and growing), 24 models now score 90% or higher, with the leader at 95.45%. Per vals.ai's own characterization, at this density a high score "no longer meaningfully distinguishes frontier models" — a density fact worth weighing regardless of how this benchmark is ultimately classified. A named, dated audit exists on the underlying question-quality question too: Epoch AI's "GPQA Diamond: What's left?" (published May 30, 2025) flagged the 40 lowest-scoring questions (of 198) for scrutiny, then deep-reviewed the 6 most extreme outliers within that set (questions where models scored under 5%), finding roughly 2.25 of those 6 invalid — extrapolating to an estimated ~8% invalid-question rate across the full set, while still concluding at least 90% of the benchmark remains valid, an error rate the piece calls comparable to what Epoch found reviewing FrontierMath.
Reading-caution on one tracked score
Claude Fable 5's reported 93.18% counts refusal-triggered fallback answers as successes; scoring refusals as failures instead drops it to 55.56% (a 41.92% refusal rate), per vals.ai, which provides a toggle to apply the correction. This is a case where the headline number and the "did the model actually answer the question" number diverge sharply.
Contamination note
GPQA's publishers explicitly ask that dataset questions not be reposted in plaintext or images online, "to reduce the risk of leakage into foundation model training corpora" — despite this, a 2025 academic contamination audit (Xu et al., "Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-Index," EMNLP 2025) flagged 4 of the 448 GPQA questions as contaminated in a Common Crawl snapshot (CC-2025-05), with at least one traced to a blog post that quoted a test-set example.
Source status (checked 2026-09-29)
Vals.ai has archived its GPQA Diamond board as saturated and no longer runs new models on it, and Artificial Analysis removed GPQA Diamond from its Intelligence Index in version 4.2 (September 2026) and — whatever its methodology page still says about running it on new releases — has published no GPQA value for any model released after 2026-09-11, Claude Opus 5.5 and Claude Sonnet 5.5 included. In practice both of this column's sources have stopped feeding new models into the table above.
Sources · 5
- Epoch AI – "GPQA Diamond: What's left?" (May 30, 2025)
- vals.ai – GPQA Diamond leaderboard
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark (arXiv:2311.12022)
- GPQA dataset card / access agreement (Hugging Face)
- Xu et al. – "Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-Index" (EMNLP 2025 / arXiv:2506.12229)
Further reading
Benchmarks related to GPQA Diamond
- Humanity's Last Exam — Reasoning, trust grade B
- HMMT Feb 2026 — Competition mathematics, trust grade C
Is the current lead on GPQA Diamond real?
TaintedGPQA Diamond is saturated — rankings here are unreliable regardless of the gap.
Written up in full, for models on this GPQA Diamond board: Claude Opus 4.8 vs DeepSeek V4 Pro (0813) · Claude Opus 4.8 vs Gemini 3.1 Pro Preview · Claude Opus 5 vs Claude Sonnet 5 — every benchmark each pair shares, not just GPQA Diamond.
GPQA Diamond scores (not ranked)
GPQA Diamond’s 26 tracked scores were each verified against their source between 2026-08-17 and 2026-09-29.
GPQA Diamond FAQ
What does GPQA Diamond measure?
PhD-level science questions written to be 'Google-proof' — hard even for skilled non-experts with internet access. The benchmark has 198 items.
How big does a gap on GPQA Diamond have to be to mean anything?
No gap on GPQA Diamond qualifies. Because the benchmark is saturated, we label every head-to-head here tainted and never apply the 7.2-point threshold its 198 items would otherwise imply.
Can you trust GPQA Diamond scores?
We grade it D and list it as saturated — frontier models now cluster near the maximum, so we label every GPQA Diamond head-to-head tainted instead of ranking its 198 items.
What is the difference between GPQA and GPQA Diamond?
Diamond is the highest-quality subset, which is not the same as the hardest. The original GPQA set has 448 graduate-level science questions; Diamond is the 198 where both expert validators answered correctly and at most one of three skilled non-experts with internet access did. That rule selects for questions that are reliably Google-proof and reliably answerable by someone who knows the field — so items the experts themselves got wrong are excluded, not included. When a leaderboard says GPQA without qualification it usually means Diamond, but the two are different sets and their scores are not interchangeable.
How many questions does GPQA Diamond have, and how is it scored?
198 questions, scored as plain accuracy on four-option multiple choice, which puts the random-guess floor at 25%. There is no LLM judge — the selected option is matched against one fixed correct answer. That simplicity is a virtue, but it is also why the benchmark saturates: with 198 items and a hard ceiling, there is very little room left between the top models.
Is GPQA Diamond still useful in 2026?
For separating frontier models, no, and this site stopped ranking it. On vals.ai's independent leaderboard, which is where the scores tracked here come from, 24 models now score 90% or higher and the leader sits at 95.45%. vals.ai, which operates that board, describes a high score at this density as no longer meaningfully distinguishing frontier models. The benchmark still measures something real; it just cannot resolve the differences people want to use it for, which is why every head-to-head here is labelled tainted rather than given a winner.
Source: https://arxiv.org/abs/2311.12022