GPQA Diamond
PhD-level science questions written to be 'Google-proof' — hard even for skilled non-experts with internet access.
How many points is real?
On GPQA Diamond, a gap smaller than 7.1 points is treated as noise — with 198 items, run-to-run variance alone can produce a difference that size. Anything bigger is worth taking seriously (subject to the caveats below).
In plain English: Grade C: Usable but small sample size or nearing saturation — treat close results with extra skepticism.
Trust caveats
Top models are clustering near the ceiling — gaps are getting smaller and noisier.
198 questions, the hardest cut of an original 448-question set. Top models now cluster within 1-3 points of each other near the ceiling — differences here are increasingly noise, not signal.
Current scores
| Gemini 3.1 Pro Preview | 95.5independent |
| GPT-5.6 Sol | 95.2independent |
| Qwen3.8-Max | 93.7independent |
| Claude Fable 5 | 93.2independent |
| Kimi K3 | 92.9independent |
Source: https://arxiv.org/abs/2311.12022