The Model Gap

LiveCodeBench

Coding contest problems dated after each model's training cutoff, so scores can't be inflated by memorization.

Items
1055
Trust grade
B
Status
current

How many points is real?

On LiveCodeBench, a gap smaller than 3.1 points is treated as noise — with 1055 items, run-to-run variance alone can produce a difference that size. Anything bigger is worth taking seriously (subject to the caveats below).

In plain English: Grade B: Reliable, mostly vendor-reported, no major known issues — but still worth an independent check.

Trust caveats

Actively discriminating between frontier models.

1,055 problems in the current release (release_v6). Grows over time by design; only compare scores from the same release version.

Current scores

Claude Fable 5
Claude Opus 5
Gemini 3.1 Pro Preview
Qwen3.8-Max
Claude Opus 4.8

Source: https://github.com/LiveCodeBench/LiveCodeBench