LiveCodeBench
Coding contest problems dated after each model's training cutoff, so scores can't be inflated by memorization.
Items
1055
Trust grade
B
Status
current
How many points is real?
On LiveCodeBench, a gap smaller than 3.1 points is treated as noise — with 1055 items, run-to-run variance alone can produce a difference that size. Anything bigger is worth taking seriously (subject to the caveats below).
In plain English: Grade B: Reliable, mostly vendor-reported, no major known issues — but still worth an independent check.
Trust caveats
Actively discriminating between frontier models.
1,055 problems in the current release (release_v6). Grows over time by design; only compare scores from the same release version.
Current scores
| Claude Fable 5 | 89.8independent |
| Claude Opus 5 | 89.0independent |
| Gemini 3.1 Pro Preview | 88.5independent |
| Qwen3.8-Max | 87.8independent |
| Claude Opus 4.8 | 87.8independent |