ARC-AGI-2

Fluid, compositional visual reasoning: inferring a hidden transformation rule from a handful of example input/output grids and applying it exactly to a new one.

Items
120
Trust grade
B
Status
current

What does an ARC-AGI-2 task look like?

One task presents a small handful of demonstration pairs — a colored grid (up to 30x30 cells, up to 10 colors) paired with the grid that results from applying a hidden transformation to it. The rule behind the pairs can be a single operation (recoloring, reflecting, resizing, counting, completing a pattern) or several rules composed and applied conditionally, and it is never stated in words — only shown. The model then sees one new, unseen input grid and must output the exact resulting grid: correct dimensions, correct colors, correct cell-by-cell layout, with no free-text reasoning credited. Grading is exact match only, so a single wrong cell fails the task; there is no partial credit and no LLM judge. No task's actual grids are reproduced here — ARC Prize keeps the scored Semi-Private and Private splits undisclosed specifically to keep this exact route out of training data, and publishing example grids on a benchmark whose whole premise is resisting memorization would work against the thing it's cited for.

How many points on ARC-AGI-2 is real?

On ARC-AGI-2, a gap smaller than 9.2 points is treated as noise — with 120 items, run-to-run variance alone can produce a difference that size. The floor itself is derived, not asserted: 100 ÷ √120 9.2, two standard errors on a benchmark this size.

Grade B: No major known issues on ARC-AGI-2, and the results here hold up — but with 120 items, close calls still deserve an independent check.

6 independent ARC-AGI-2 scores — dots inside one dashed band are closer than ARC-AGI-2’s 9.2-point noise floor, statistically indistinguishable.

Trust caveats

ARC-AGI-2 is actively discriminating between frontier models — a gap of at least 9.2 points here is worth taking seriously, though it only counts as a confirmed real gap once both scores come from independent runs.

Dataset and maintainer

ARC-AGI-2 is built and maintained by ARC Prize (François Chollet, Mike Knoop, Gregory Kamradt, Bryan Landers, Henry Pinkard), announced 2025-03-24 and detailed in "ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems" (arXiv:2505.11831, submitted 2025-05-17). It has four splits — Training (1,000, public), Public Eval (120, public), Semi-Private Eval (120, undisclosed, official leaderboard set), Private Eval (120, undisclosed, Kaggle final-contest set) — and the scores tracked on this page are drawn from the 120-task Semi-Private Eval Set, per arcprize.org/arc-agi/2's Dataset Structure table.

Scoring mechanics

Each task is graded by exact match against a single correct output grid — no partial credit for a near-miss, no LLM judge, no multiple-choice guess floor (0% is possible). Difficulty is calibrated against measured human performance rather than set by intuition: ARC Prize ran a controlled study with 400+ participants and kept only tasks "solved pass@2 by at least two humans" (arcprize.org/blog/announcing-arc-agi-2-and-arc-prize-2025, 2025-03-24) — the 100% human-solvable ceiling this page's scores are read against comes directly from that calibration criterion, not an assumption about human ability.

Reasoning-effort-tier caveat

A bare ARC-AGI-2 percentage is close to meaningless without its reasoning/compute tier attached, because the tier swings scores by tens of points independent of model generation — and sometimes the higher tier's cell is simply blank rather than populated at all. This page's own score table shows both patterns on the official leaderboard: Claude Opus 4.8's Max-tier cell is genuinely N/A for ARC-AGI-2, so its usable ceiling is its High tier at 72.1%, while Claude Opus 5 — one generation newer — has both tiers populated, at 88.3% (High) and 90.4% (Max) (arcprize.org/leaderboard, per this page's own score rows). ARC Prize's own leaderboard is built around this: its "Reasoning Systems Trend Line" plots "the same model at different reasoning levels" as connected points precisely because one number understates the spread (arcprize.org/leaderboard). Every score on this page has to be read together with its variant tag for that reason.

Saturation/ceiling evidence

This benchmark sits meaningfully below its ceiling, not at it. As of 2026-08-22 the top tracked score is GPT-5.6 Sol at 92.5%, with Claude Opus 5 at 90.4% (benchlm.ai/benchmarks/arc-agi-2) — up sharply from a ~83-85% top cluster in February-March 2026 — Gemini 3 Deep Think 84.6% (2026-02-12) and GPT-5.4 Pro (XHigh) 83.3% (2026-03-04), both per ARC Prize's own leaderboard — but still short of the 100% human-panel ceiling above. ARC Prize built the benchmark specifically to resist the saturation pattern ARC-AGI-1 hit: its own framing states "log-linear scaling is insufficient to beat ARC-AGI-2" and that new test-time adaptation methods, not scale, are what closes the remaining gap (arcprize.org/arc-agi/2) — and the organization is already running an ARC-AGI-3 competition as the next tier, rather than treating this one as a permanent ceiling.

Anti-contamination design

Instead of a license clause asking people not to republish questions, ARC-AGI-2's protection is structural. The scored Semi-Private Eval Set is "not public" and may at most have "been exposed to limited third-parties (e.g., via API)"; the Private Eval Set used for the final Kaggle ranking "has not been exposed to third-parties" at all (arcprize.org/blog/announcing-arc-agi-2-and-arc-prize-2025, 2025-03-24) — so there is no public copy of the graded items for a training run to have absorbed. This distinction is not just theoretical: Imbue's own February 2026 writeup measured its 88.1%-to-95.1% improvement on the Public Eval Set and explicitly flags that the result is "not 100% comparable" to Semi-Private leaderboard scores, precisely because the public set is the one anyone can inspect (imbue.com/blog/2026-02-27-arc-agi-2-evolution). Checked arcprize.org directly (2026-08-24): no explicit copyright or republication-restriction clause covers individual task grids, so the safeguard here is the undisclosed split, not a legal one — which is also why this page's task_shape describes the format only and reproduces no item.

Trust grade

B. Every material claim above — maintainer, calibration method, scoring mechanics, tier sensitivity, saturation trend — has a named, dated source, and the anti-contamination design is structural rather than a disclosure clause nobody checks. It stops short of this site's A grade (Agents' Last Exam) because ARC-AGI-2's Semi-Private pool is fixed until the next major version rather than refreshed on a rolling schedule, so a long-lived leak into that fixed pool would not self-heal the way a rolling-refresh benchmark's does.

Is the current lead on ARC-AGI-2 real?

TieA 2.1-point gap on ARC-AGI-2 (n=120) is inside the noise band (±9.2) — call it a tie.

Computed on the max variant, the same one charted above — other variants of ARC-AGI-2 can rank differently and are shown separately in the table below.

Current scores

ARC-AGI-2’s 11 tracked scores were verified against their sources on or after 2026-08-24.

max
GPT-5.6 Sol
Claude Opus 5
Claude Fable 5
DeepSeek V4 Flash (0731)
DeepSeek V4 Pro (0813)
Kimi K3
high
Gemini 3.7 Flash
Claude Opus 4.8
xhigh
Grok 4.6
untiered
Gemini 3.1 Pro Preview
GLM-5.2

FAQ

What does ARC-AGI-2 measure?

Fluid, compositional visual reasoning: inferring a hidden transformation rule from a handful of example input/output grids and applying it exactly to a new one. The benchmark has 120 items.

How big does a gap on ARC-AGI-2 have to be to mean anything?

At least 9.2 points. Across 120 items, a smaller gap sits inside the spread that repeat runs produce on their own, so we call it a tie — provided both scores are independent runs; if either side is a vendor claim, the row is labelled unverified instead.

Can you trust ARC-AGI-2 scores?

We grade it B and list it as current — it still tells frontier models apart across 120 items. A gap of at least 9.2 points clears that noise floor, but it only counts as a real gap once both scores come from independent runs.

Source: https://arcprize.org/leaderboard