LiveBench

LiveBench's tracked "Overall" score is a composite average across seven independently-graded categories — reasoning, coding, agentic coding, math, data analysis, language, and instruction following — each scored against an objective, verifiable ground-truth answer rather than any single skill.

Items
1436
Trust grade
A
Status
current

What does a LiveBench task look like?

LiveBench is unlike every other benchmark tracked on this site in one specific way: the tracked "Overall" score isn't built from one uniform question type, because it averages seven separately-run categories, each with its own task format. Reasoning and Language items are closed-ended questions or short-answer puzzles graded against a single fixed correct answer. Math items are drawn from recent competition-style problems with a numeric or symbolic final answer. Coding items ask for a single-turn program that must pass hidden held-out test cases, while Agentic Coding items pose a similar goal but require the model to work iteratively inside a tool-using environment rather than emit one finished answer in one shot. Data Analysis items hand the model a real or synthetic dataset and ask it to compute or transform something checkable against a known result. Instruction Following items test whether a response obeys a set of stated constraints — length, format, required or forbidden content — that can be verified programmatically rather than judged by feel. Across all seven categories, source material is deliberately current: recent arXiv papers, news articles, Kaggle/Socrata datasets, IMDb synopses, and newly-published math-competition problems, refreshed periodically so older, possibly-memorized questions get retired. No example item is reproduced here — only the general shape of each category is described.

How many points on LiveBench is real?

On LiveBench, a gap smaller than 2.7 points is treated as noise — with 1436 items, run-to-run variance alone can produce a difference that size. The floor itself is derived, not asserted: 100 ÷ √1436 2.7, two standard errors on a benchmark this size.

Grade A: LiveBench is designed to resist contamination (e.g. refreshed on a schedule). With 1436 items, trust close results here.

13 independent LiveBench scores — dots inside one dashed band are closer than LiveBench’s 2.7-point noise floor, statistically indistinguishable.

Trust caveats

LiveBench is actively discriminating between frontier models — a gap of at least 2.7 points here is worth taking seriously, though it only counts as a confirmed real gap once both scores come from independent runs.

Methodology

LiveBench is run entirely on the maintainers' own harness — every score on this page comes from the LiveBench team's own evaluation code (github.com/livebench/livebench), not an aggregated vendor self-report or a third party's independent re-scraping of vendor claims. Each model gets exactly one row run under one fixed configuration — LiveBench doesn't offer a selectable reasoning-effort menu the way some other benchmarks tracked on this site do, so a row's label mentioning "Max Effort" or a tier name describes what configuration that run used, not one option among several — every tracked score here shares a single 'default' variant and compares directly against every other one. Grading is objective and automated in every one of the seven categories: each question resolves to a single verifiable ground truth (a fixed answer, a specific numeric result, a passing hidden test suite, or a programmatically-checkable constraint), so there is no LLM-judge step deciding whether an answer "looks right." The question pool is refreshed periodically with material sourced from recent arXiv papers, news articles, Kaggle/Socrata datasets, IMDb synopses, and newly-published math-competition problems, and older questions are retired over time — the contamination-resistance property the benchmark was purpose-built around (Colin White et al., "LiveBench: A Challenging, Contamination-Limited LLM Benchmark," arXiv:2406.19314, ICLR 2025 Spotlight).

Item-count caveat

This site could not find a reliable current item count for the 2026-06-25 release despite checking four independent sources — the live leaderboard states "23 tasks" with no item count given, the GitHub README is stale at "18 tasks / 6 categories," the official Hugging Face dataset repos (livebench/reasoning, /coding, /math, /data_analysis, /language, /instruction_following) are stale at an April 2025 snapshot and cover only 6 of the current 7 categories, and the arXiv paper's latest revision (v2, April 2025) predates the current release. The meaningful_gap figure below (2.7) is therefore built on that stale Hugging Face count — 1,436 items summed across those 6 category repos — used deliberately as a conservative lower-bound estimate, not a measured figure for the current release. It is known incomplete: it omits the 7th category (Agentic Coding) entirely, so the true current item count is higher, and the true statistical noise floor is probably somewhat tighter than 2.7 implies.

Saturation

The tracked "Overall" score is not saturated — across 43 ranked models on the current release, scores span roughly 62.3 to 83.0, well short of a ceiling. That headline range hides real category-level clustering, though: among the top ~10 models, Data Analysis subscores mostly sit in the 90-96% band and Reasoning subscores mostly sit in the 87-92% band, both closer to saturated than the Overall figure suggests. Coding, and especially Agentic Coding, remain far from that ceiling — the top Agentic Coding subscores are still only in the 60s. This site classifies the tracked score as "current" rather than "saturated" on that basis, but a model comparison that looks tight on Overall may be built from components with very different amounts of remaining headroom.

Composite caveat

"Overall" is an unweighted average across seven categories that measure genuinely different things — closed-ended reasoning and language questions, competition math, single-turn coding, iterative tool-using agentic coding, dataset-driven data analysis, and constraint-following. Two models can land on the same Overall number by different routes — one strong in Agentic Coding and weak in Math, another the reverse — and the average erases that difference entirely. This site doesn't currently surface LiveBench's per-category breakdown, so a reader treating a tie (or a narrow gap) on this page as "these two models perform equivalently" should know that conclusion only holds at the composite level. This averaging effect is exactly why LiveBench needed its own new benchmark category here (composite/general-knowledge) rather than being forced into this site's execution-form buckets, and why this score should inform, not replace, the category-specific benchmarks tracked elsewhere on this site.

Trust grade

A. LiveBench is run entirely on the maintainers' own harness with objective, automated grading and no LLM-judge step, is peer-reviewed (ICLR 2025 Spotlight), and is purpose-built with a periodic-refresh mechanism that retires old questions before they can be memorized — the same caliber of evidence behind this site's other A grade (Agents' Last Exam). The one asterisk is bookkeeping, not methodology: the site cannot currently confirm the current release's exact item count (see the caveat above), which affects how precisely meaningful_gap is calibrated but not whether the benchmark's own grading is trustworthy.

Is the current lead on LiveBench real?

TieA 2.0-point gap on LiveBench (n=1436) is inside the noise band (±2.7) — call it a tie.

Current scores

LiveBench’s 13 tracked scores were verified against their sources on or after 2026-08-24.

Claude Fable 5
GPT-5.6 Sol
Claude Opus 5
Kimi K3
Gemini 3.7 Flash
Qwen3.8-Max
Grok 4.6
DeepSeek V4 Pro (0813)
Gemini 3.1 Pro Preview
Claude Opus 4.8
Claude Sonnet 5
DeepSeek V4 Flash (0731)
GLM-5.2

FAQ

What does LiveBench measure?

LiveBench's tracked "Overall" score is a composite average across seven independently-graded categories — reasoning, coding, agentic coding, math, data analysis, language, and instruction following — each scored against an objective, verifiable ground-truth answer rather than any single skill. The benchmark has 1436 items.

How big does a gap on LiveBench have to be to mean anything?

At least 2.7 points. Across 1436 items, a smaller gap sits inside the spread that repeat runs produce on their own, so we call it a tie — provided both scores are independent runs; if either side is a vendor claim, the row is labelled unverified instead.

Can you trust LiveBench scores?

We grade it A and list it as current — it still tells frontier models apart across 1436 items. A gap of at least 2.7 points clears that noise floor, but it only counts as a real gap once both scores come from independent runs.

Source: https://livebench.ai/