LiveBench

LiveBench's tracked "Overall" score is a composite average across seven independently-graded categories — reasoning, coding, agentic coding, math, data analysis, language, and instruction following — each scored against an objective, verifiable ground-truth answer rather than any single skill.

Items
1436
Trust grade
A
Status
current

What does a LiveBench task look like?

LiveBench is unlike every other benchmark tracked on this site in one specific way: the tracked "Overall" score isn't built from one uniform question type, because it averages seven separately-run categories, each with its own task format. Reasoning and Language items are closed-ended questions or short-answer puzzles graded against a single fixed correct answer. Math items are drawn from recent competition-style problems with a numeric or symbolic final answer. Coding items ask for a single-turn program that must pass hidden held-out test cases, while Agentic Coding items pose a similar goal but require the model to work iteratively inside a tool-using environment rather than emit one finished answer in one shot. Data Analysis items hand the model a real or synthetic dataset and ask it to compute or transform something checkable against a known result. Instruction Following items test whether a response obeys a set of stated constraints — length, format, required or forbidden content — that can be verified programmatically rather than judged by feel. Across all seven categories, source material is deliberately current: recent arXiv papers, news articles, Kaggle/Socrata datasets, IMDb synopses, and newly-published math-competition problems, refreshed periodically so older, possibly-memorized questions get retired. No example item is reproduced here — only the general shape of each category is described.

How many points on LiveBench is real?

On LiveBench, a gap smaller than 2.7 points is treated as noise — with 1436 items, run-to-run variance alone can produce a difference that size. The floor itself is derived, not asserted: 100 ÷ √1436 ≈ 2.7, two standard errors on LiveBench at this size.

Grade A: LiveBench is designed to resist contamination (e.g. refreshed on a schedule). With 1436 items, trust close results here.

Can you trust LiveBench? Caveats

LiveBench is actively discriminating between frontier models — a gap of at least 2.7 points here is worth taking seriously, though it only counts as a confirmed real gap once both scores come from independent runs.

Methodology

LiveBench is run entirely on the maintainers' own harness — every score on this page comes from the LiveBench team's own evaluation code (github.com/livebench/livebench), not an aggregated vendor self-report or a third party's independent re-scraping of vendor claims. A row's label naming a tier ("Max Effort", "xHigh Effort") describes the configuration that run used. Since September 2026 the board carries separate effort-tier rows for some models (Claude Opus 5.5 and Claude Sonnet 5.5 among them, checked 2026-09-29), showing the highest-scoring one by default and the rest behind a variants toggle; this site records the default-displayed row, names the tier in each score's notes, and keeps a single 'default' variant, so every tracked score compares directly against every other one — a tier mismatch between two models is disclosed in prose, not blocked by the harness gate. Grading is objective and automated in every one of the seven categories: each question resolves to a single verifiable ground truth (a fixed answer, a specific numeric result, a passing hidden test suite, or a programmatically-checkable constraint), so there is no LLM-judge step deciding whether an answer "looks right." The question pool is refreshed periodically with material sourced from recent arXiv papers, news articles, Kaggle/Socrata datasets, IMDb synopses, and newly-published math-competition problems, and older questions are retired over time — the contamination-resistance property the benchmark was purpose-built around (Colin White et al., "LiveBench: A Challenging, Contamination-Limited LLM Benchmark," arXiv:2406.19314, ICLR 2025 Spotlight).

Item-count caveat

This site could not find a reliable current item count for the 2026-06-25 release despite checking four independent sources — the live leaderboard states "23 tasks" with no item count given, the GitHub README is stale at "18 tasks / 6 categories," the official Hugging Face dataset repos (livebench/reasoning, /coding, /math, /data_analysis, /language, /instruction_following) are stale at an April 2025 snapshot and cover only 6 of the current 7 categories, and the arXiv paper's latest revision (v2, April 2025) predates the current release. The meaningful_gap figure below (2.7) is therefore built on that stale Hugging Face count — 1,436 items summed across those 6 category repos — used deliberately as a conservative lower-bound estimate, not a measured figure for the current release. It is known incomplete: it omits the 7th category (Agentic Coding) entirely, so the true current item count is higher, and the true statistical noise band is probably somewhat tighter than 2.7 implies.

Saturation

The tracked "Overall" score is not saturated — across 43 ranked models on the current release, scores span roughly 62.3 to 83.0, well short of a ceiling. That headline range hides real category-level clustering, though: among the top ~10 models, Data Analysis subscores mostly sit in the 90-96% band and Reasoning subscores mostly sit in the 87-92% band, both closer to saturated than the Overall figure suggests. Coding, and especially Agentic Coding, remain far from that ceiling — the top Agentic Coding subscores are still only in the 60s. This site classifies the tracked score as "current" rather than "saturated" on that basis, but a model comparison that looks tight on Overall may be built from components with very different amounts of remaining headroom.

Composite caveat

"Overall" is an unweighted average across seven categories that measure genuinely different things — closed-ended reasoning and language questions, competition math, single-turn coding, iterative tool-using agentic coding, dataset-driven data analysis, and constraint-following. Two models can land on the same Overall number by different routes — one strong in Agentic Coding and weak in Math, another the reverse — and the average erases that difference entirely. This site doesn't currently surface LiveBench's per-category breakdown, so a reader treating a tie (or a narrow gap) on this page as "these two models perform equivalently" should know that conclusion only holds at the composite level. This averaging effect is exactly why LiveBench needed its own new benchmark category here (composite/general-knowledge) rather than being forced into this site's execution-form buckets, and why this score should inform, not replace, the category-specific benchmarks tracked elsewhere on this site.

Trust grade

A. LiveBench is run entirely on the maintainers' own harness with objective, automated grading and no LLM-judge step, is peer-reviewed (ICLR 2025 Spotlight), and is purpose-built with a periodic-refresh mechanism that retires old questions before they can be memorized — the same caliber of evidence behind this site's other A grade (Agents' Last Exam). The one asterisk is bookkeeping, not methodology: the site cannot currently confirm the current release's exact item count (see the caveat above), which affects how precisely meaningful_gap is calibrated but not whether the benchmark's own grading is trustworthy.

Is the current lead on LiveBench real?

TieA 0.2-point gap on LiveBench (n=1436) is inside the noise band (±2.7) — call it a tie.

Written up in full, for models on this LiveBench board: Claude Opus 5 vs Claude Sonnet 5 · Gemini 3.1 Pro Preview vs Kimi K3 · Claude Opus 4.8 vs DeepSeek V4 Pro (0813) — every benchmark each pair shares, not just LiveBench.

LiveBench leaderboard: current scores

LiveBench’s 30 tracked scores were each verified against their source between 2026-08-24 and 2026-10-08.

LiveBench scores
ModelLiveBench score
Claude Fable 5.1
Claude Opus 5.5
Claude Fable 5
Muse Spark 1.3
GPT-6.1 Sol
DeepSeek V4.1 Flash
GPT-5.6 Sol
Claude Opus 5
GPT-6 Sol
Kimi K3
Gemini 3.7 Flash
Qwen3.8-Max
Grok 4.6
Muse Spark 1.2
Claude Sonnet 5.5
DeepSeek V4 Pro (0813)
Grok 4.7
Gemini 3.1 Pro Preview
Claude Opus 4.8
Qwen3.8-Flash-Next
GLM-5.3
Claude Sonnet 5
Gemini 3.8 Flash
DeepSeek V4 Flash (0731)
GPT-5.6 Luna
GLM-5.2
Claude Haiku 5.5
GPT-6 Luna
GLM-5.3-Flash
MiniMax M3

LiveBench FAQ

What does LiveBench measure?

LiveBench's tracked "Overall" score is a composite average across seven independently-graded categories — reasoning, coding, agentic coding, math, data analysis, language, and instruction following — each scored against an objective, verifiable ground-truth answer rather than any single skill. LiveBench has 1436 items.

How big does a gap on LiveBench have to be to mean anything?

At least 2.7 points on LiveBench. Across LiveBench's 1436 items, a smaller gap sits inside the spread that repeat runs produce on their own, so we call it a tie — provided both scores are independent runs; if either side is a vendor claim, the row is labelled unverified instead.

Can you trust LiveBench scores?

We grade it A and list it as current — it still tells frontier models apart across 1436 LiveBench items. A gap of at least 2.7 points clears LiveBench's own noise band, but it only counts as a real gap once both scores come from independent runs.

Who runs LiveBench scores?

The LiveBench team, on their own harness — every score on this page comes from their evaluation code (github.com/livebench/livebench), not from vendors self-reporting or a third party re-scraping claims. A row's effort-tier label ("Max Effort", "xHigh Effort") describes the configuration that particular run used.

Does LiveBench use an LLM judge?

No. Each of the seven categories resolves to a single verifiable ground truth — a fixed answer, a specific numeric result, a passing hidden test suite, or a programmatically-checkable constraint — so no judge model ever decides whether an answer "looks right." The question pool is also refreshed periodically from recent arXiv papers, news articles, Kaggle/Socrata datasets, IMDb synopses, and newly-published math competitions to resist memorization.

What are LiveBench's seven categories?

Reasoning, coding, agentic coding, math, data analysis, language, and instruction following. The score tracked on this page is the Overall composite — an average across those seven categories — which is why this site files LiveBench as a composite benchmark rather than a single-skill one.

Benchmarks related to LiveBench

Source: https://livebench.ai/