The Model Gap

GPQA Diamond

PhD-level science questions written to be 'Google-proof' — hard even for skilled non-experts with internet access.

Items
198
Trust grade
C
Status
aging

How many points is real?

On GPQA Diamond, a gap smaller than 7.1 points is treated as noise — with 198 items, run-to-run variance alone can produce a difference that size. Anything bigger is worth taking seriously (subject to the caveats below).

In plain English: Grade C: Usable but small sample size or nearing saturation — treat close results with extra skepticism.

Trust caveats

Top models are clustering near the ceiling — gaps are getting smaller and noisier.

198 questions, the hardest cut of an original 448-question set. Top models now cluster within 1-3 points of each other near the ceiling — differences here are increasingly noise, not signal.

Current scores

Gemini 3.1 Pro Preview
GPT-5.6 Sol
Qwen3.8-Max
Claude Fable 5
Kimi K3

Source: https://arxiv.org/abs/2311.12022