LiveCodeBench

Coding contest problems dated after each model's training cutoff, so scores can't be inflated by memorization.

Items
1055
Trust grade
D
Status
saturated

What does a LiveCodeBench task look like?

One task is a single competitive-programming problem pulled from a real contest on LeetCode, AtCoder, or Codeforces, tagged with the date it was published. The model is given the natural-language problem statement — the scenario, constraints, and required input/output format — and must produce a complete program or function that solves it. Grading is fully automated and execution-based: the submitted code is run against a hidden suite of test cases under a time and memory limit, using an APPS-style checker, and the problem counts as solved only if every hidden test passes. No human or model judge is involved, and there is no multiple-choice floor since the model must generate original code.

How many points on LiveCodeBench is real?

On LiveCodeBench the 3.1-point noise band is beside the point: the benchmark is saturated, so every comparison across its 1055 items is labeled tainted instead of being measured against a threshold.

Grade D: LiveCodeBench is saturated or has known contamination issues — its rankings are unreliable at any gap size, 1055 items or not.

17 independent LiveCodeBench scores, no noise band drawn: the benchmark is saturated, so every LiveCodeBench head-to-head is labelled tainted rather than measured against its 3.1-point threshold. How tightly the LiveCodeBenchdots bunch above is the reason for that call.

Can you trust LiveCodeBench? Caveats

Most frontier models score near the maximum on LiveCodeBench — this benchmark no longer separates them well.

1,055 problems in the current release (release_v6), spanning problems released May 2023–April 2025 (per the official repo's dataset-versions table). Grows over time by design; only compare scores from the same release version. The benchmark started at 400 problems in v1 (May 2023–March 2024) and has been extended release by release since — v6 is roughly 2.6x the size of v1. Created and maintained by researchers at UC Berkeley's Sky Computing Lab, MIT, and Cornell (Naman Jain, Ion Stoica, Armando Solar-Lezama, Koushik Sen and co-authors), published at ICLR 2025.

Contamination resistance is structural, not filter-based

Every problem carries a release date, and the intended practice is to score a model only on problems released after its training cutoff, so a correct solve can't be explained by the problem having leaked into pretraining data. The paper frames this as a risk it is designed to close rather than a confirmed flaw in prior work: it states that existing benchmarks like HumanEval, MBPP, and APPS "may be subject to potential contamination or overfitting" because problem samples can appear in pretraining corpora, and backs the concern with a direct measurement on this benchmark's own time-stamped problems — DeepSeek-Instruct-33B's accuracy drops sharply on problems dated just before its release, consistent with earlier problems having leaked into its training data. Note that the aggregate release_v6 scores tracked here are not necessarily each model's own cutoff-filtered subset — that distinction matters when comparing models with different training cutoffs on the same release.

Scoring is pass@1

A model's generated code is executed against a hidden test suite for each problem, using a modified version of the checker from the APPS benchmark (the official repo notes they fixed unhandled edge cases in the original APPS checker); a problem counts as solved only if the submission passes every hidden test within the time/memory limit. pass@5 is also reported by the official harness. There is no LLM judge and no chance floor — this is free-form code generation, not multiple choice.

Trust-grade caveat

The problem-sourcing design is contamination-resistant, but the reference evaluation harness has a documented reliability gap. Collinear AI published a dated audit (Aug 12, 2025, "Leveling the Playing Field: LiveCodeBench's Big Bug Fix") showing the official scoring script could truncate correct model output at a spurious "###" stop token, misidentify non-code text inside backticks as the submitted answer, and apply hard-coded rather than API-native chat templates — together swinging measured scores by roughly 50% relative for at least one model in their reproduction (their Qwen3-8B run initially read 38.3% vs. the model's official technical-report figure of ~57.8%, converging to ~57.8% only after they patched the harness). The corresponding fix PRs (#117 "Default stop flag to None" and #118 "Added precise python backtick check") were opened against the official LiveCodeBench repo the same day and remained open/unmerged as of this review, so scores collected on different harness versions or configurations may not be directly comparable even within the same release_v6 problem set.

Saturation call (2026-09-29)

Graded D and listed as saturated, the treatment already applied to SWE-bench Verified and GPQA Diamond. Two pieces of evidence. First, vals.ai — the evaluator behind every score on this page — has stopped running the benchmark on new models, on the stated ground that performance has saturated. Second, the tracked scores themselves: 12 of the 17 independent rows sit within 3.33 points of the leader (90.52 down to 87.19), so a single 3.1-point noise band holds most of the frontier, and the bottom of the field is separated mainly by age (the lowest row, GLM-5.2 at 69.5, is a superseded model). Under this call compareScores() returns Tainted for every LiveCodeBench pair — the 136 comparable pairs that previously read 66 real gaps and 70 ties — because a close score on a ceiling-bound test says the test has run out of headroom, not that two models are matched. The contamination-resistant design described above is unchanged; what has changed is that frontier models now clear the current release set.

Source status (checked 2026-09-29)

Vals.ai, the only evaluator behind this column, now shows an "Archived Benchmark" banner on its LiveCodeBench board — "Since performance on this benchmark has saturated, we no longer run this benchmark on new model releases" (page dated 9/1/2026, though it still scored Gemini 3.8 Flash, released 2026-09-02). Models released after early September 2026 will not gain a row here from that source.

Benchmarks related to LiveCodeBench

Is the current lead on LiveCodeBench real?

TaintedLiveCodeBench is saturated — rankings here are unreliable regardless of the gap.

Written up in full, for models on this LiveCodeBench board: Claude Opus 4.8 vs DeepSeek V4 Pro (0813) · Claude Opus 4.8 vs Gemini 3.1 Pro Preview · Claude Opus 5 vs Claude Sonnet 5 — every benchmark each pair shares, not just LiveCodeBench.

LiveCodeBench scores (not ranked)

LiveCodeBench’s 17 tracked scores were each verified against their source between 2026-08-17 and 2026-09-03.

LiveCodeBench scores
ModelLiveCodeBench score
Claude Fable 5.1
Claude Fable 5
Gemini 3.8 Flash
Claude Opus 5
Gemini 3.7 Flash
Gemini 3.1 Pro Preview
Grok 4.6
Qwen3.8-Max
Claude Opus 4.8
DeepSeek V4 Pro (0813)
DeepSeek V4 Flash (0731)
Kimi K3
GPT-5.6 Sol
Claude Sonnet 5
MiniMax M3
GLM-5.3
GLM-5.2

LiveCodeBench FAQ

What does LiveCodeBench measure?

Coding contest problems dated after each model's training cutoff, so scores can't be inflated by memorization. The benchmark has 1055 items.

How big does a gap on LiveCodeBench have to be to mean anything?

No gap on LiveCodeBench qualifies. Because the benchmark is saturated, we label every head-to-head here tainted and never apply the 3.1-point threshold its 1055 items would otherwise imply.

Can you trust LiveCodeBench scores?

We grade it D and list it as saturated — frontier models now cluster near the maximum, so we label every LiveCodeBench head-to-head tainted instead of ranking its 1055 items.

Source: https://github.com/LiveCodeBench/LiveCodeBench