LiveCodeBench
Coding contest problems dated after each model's training cutoff, so scores can't be inflated by memorization.
What does a LiveCodeBench task look like?
One task is a single competitive-programming problem pulled from a real contest on LeetCode, AtCoder, or Codeforces, tagged with the date it was published. The model is given the natural-language problem statement — the scenario, constraints, and required input/output format — and must produce a complete program or function that solves it. Grading is fully automated and execution-based: the submitted code is run against a hidden suite of test cases under a time and memory limit, using an APPS-style checker, and the problem counts as solved only if every hidden test passes. No human or model judge is involved, and there is no multiple-choice floor since the model must generate original code.
How many points on LiveCodeBench is real?
On LiveCodeBench the 3.1-point noise band is beside the point: the benchmark is saturated, so every comparison across its 1055 items is labeled tainted instead of being measured against a threshold.
Grade D: LiveCodeBench is saturated or has known contamination issues — its rankings are unreliable at any gap size, 1055 items or not.
17 independent LiveCodeBench scores, no noise band drawn: the benchmark is saturated, so every LiveCodeBench head-to-head is labelled tainted rather than measured against its 3.1-point threshold. How tightly the LiveCodeBenchdots bunch above is the reason for that call.
Can you trust LiveCodeBench? Caveats
Most frontier models score near the maximum on LiveCodeBench — this benchmark no longer separates them well.
1,055 problems in the current release (release_v6), spanning problems released May 2023–April 2025 (per the official repo's dataset-versions table). Grows over time by design; only compare scores from the same release version. The benchmark started at 400 problems in v1 (May 2023–March 2024) and has been extended release by release since — v6 is roughly 2.6x the size of v1. Created and maintained by researchers at UC Berkeley's Sky Computing Lab, MIT, and Cornell (Naman Jain, Ion Stoica, Armando Solar-Lezama, Koushik Sen and co-authors), published at ICLR 2025.
Contamination resistance is structural, not filter-based
Every problem carries a release date, and the intended practice is to score a model only on problems released after its training cutoff, so a correct solve can't be explained by the problem having leaked into pretraining data. The paper frames this as a risk it is designed to close rather than a confirmed flaw in prior work: it states that existing benchmarks like HumanEval, MBPP, and APPS "may be subject to potential contamination or overfitting" because problem samples can appear in pretraining corpora, and backs the concern with a direct measurement on this benchmark's own time-stamped problems — DeepSeek-Instruct-33B's accuracy drops sharply on problems dated just before its release, consistent with earlier problems having leaked into its training data. Note that the aggregate release_v6 scores tracked here are not necessarily each model's own cutoff-filtered subset — that distinction matters when comparing models with different training cutoffs on the same release.
Scoring is pass@1
A model's generated code is executed against a hidden test suite for each problem, using a modified version of the checker from the APPS benchmark (the official repo notes they fixed unhandled edge cases in the original APPS checker); a problem counts as solved only if the submission passes every hidden test within the time/memory limit. pass@5 is also reported by the official harness. There is no LLM judge and no chance floor — this is free-form code generation, not multiple choice.
Trust-grade caveat
The problem-sourcing design is contamination-resistant, but the reference evaluation harness has a documented reliability gap. Collinear AI published a dated audit (Aug 12, 2025, "Leveling the Playing Field: LiveCodeBench's Big Bug Fix") showing the official scoring script could truncate correct model output at a spurious "###" stop token, misidentify non-code text inside backticks as the submitted answer, and apply hard-coded rather than API-native chat templates — together swinging measured scores by roughly 50% relative for at least one model in their reproduction (their Qwen3-8B run initially read 38.3% vs. the model's official technical-report figure of ~57.8%, converging to ~57.8% only after they patched the harness). The corresponding fix PRs (#117 "Default stop flag to None" and #118 "Added precise python backtick check") were opened against the official LiveCodeBench repo the same day and remained open/unmerged as of this review, so scores collected on different harness versions or configurations may not be directly comparable even within the same release_v6 problem set.
Saturation call (2026-09-29)
Graded D and listed as saturated, the treatment already applied to SWE-bench Verified and GPQA Diamond. Two pieces of evidence. First, vals.ai — the evaluator behind every score on this page — has stopped running the benchmark on new models, on the stated ground that performance has saturated. Second, the tracked scores themselves: 12 of the 17 independent rows sit within 3.33 points of the leader (90.52 down to 87.19), so a single 3.1-point noise band holds most of the frontier, and the bottom of the field is separated mainly by age (the lowest row, GLM-5.2 at 69.5, is a superseded model). Under this call compareScores() returns Tainted for every LiveCodeBench pair — the 136 comparable pairs that previously read 66 real gaps and 70 ties — because a close score on a ceiling-bound test says the test has run out of headroom, not that two models are matched. The contamination-resistant design described above is unchanged; what has changed is that frontier models now clear the current release set.
Source status (checked 2026-09-29)
Vals.ai, the only evaluator behind this column, now shows an "Archived Benchmark" banner on its LiveCodeBench board — "Since performance on this benchmark has saturated, we no longer run this benchmark on new model releases" (page dated 9/1/2026, though it still scored Gemini 3.8 Flash, released 2026-09-02). Models released after early September 2026 will not gain a row here from that source.
Sources · 6
- LiveCodeBench official GitHub README (dataset versions, checker methodology)
- LiveCodeBench paper, Jain et al., ICLR 2025 (authors, contamination-risk framing, DeepSeek/GPT-4-O contamination evidence)
- UC Berkeley Sky Computing Lab project page
- Collinear AI, "Leveling the Playing Field: LiveCodeBench's Big Bug Fix" (Aug 12, 2025)
- LiveCodeBench GitHub PR #117, "Default stop flag to None" (opened Aug 12, 2025, open/unmerged as of review)
- LiveCodeBench GitHub PR #118, "Added precise python backtick check" (opened Aug 12, 2025, open/unmerged as of review)
Benchmarks related to LiveCodeBench
- SWE-bench Verified — Bug fixing, trust grade D
- DeepSWE — Long-horizon coding, trust grade B
Is the current lead on LiveCodeBench real?
TaintedLiveCodeBench is saturated — rankings here are unreliable regardless of the gap.
Written up in full, for models on this LiveCodeBench board: Claude Opus 4.8 vs DeepSeek V4 Pro (0813) · Claude Opus 4.8 vs Gemini 3.1 Pro Preview · Claude Opus 5 vs Claude Sonnet 5 — every benchmark each pair shares, not just LiveCodeBench.
LiveCodeBench scores (not ranked)
LiveCodeBench’s 17 tracked scores were each verified against their source between 2026-08-17 and 2026-09-03.
| Model | LiveCodeBench score |
|---|---|
| Claude Fable 5.1 | 90.5independent |
| Claude Fable 5 | 89.8independent |
| Gemini 3.8 Flash | 89.5independent |
| Claude Opus 5 | 89.0independent |
| Gemini 3.7 Flash | 88.7independent |
| Gemini 3.1 Pro Preview | 88.5independent |
| Grok 4.6 | 88.2independent |
| Qwen3.8-Max | 87.8independent |
| Claude Opus 4.8 | 87.8independent |
| DeepSeek V4 Pro (0813) | 87.5independent |
| DeepSeek V4 Flash (0731) | 87.3independent |
| Kimi K3 | 87.2independent |
| GPT-5.6 Sol | 82.6independent |
| Claude Sonnet 5 | 82.4independent |
| MiniMax M3 | 82.2independent |
| GLM-5.3 | 80.5independent |
| GLM-5.2 | 69.5independent |
LiveCodeBench FAQ
What does LiveCodeBench measure?
Coding contest problems dated after each model's training cutoff, so scores can't be inflated by memorization. The benchmark has 1055 items.
How big does a gap on LiveCodeBench have to be to mean anything?
No gap on LiveCodeBench qualifies. Because the benchmark is saturated, we label every head-to-head here tainted and never apply the 3.1-point threshold its 1055 items would otherwise imply.
Can you trust LiveCodeBench scores?
We grade it D and list it as saturated — frontier models now cluster near the maximum, so we label every LiveCodeBench head-to-head tainted instead of ranking its 1055 items.