Compare AI models
Choose anywhere from two to five models and every benchmark score and listed price we have sourced for them lines up in one table.
Narrow it down to exactly two and the table adds a Signal column — the same real-gap-or-noise call this site runs on every head-to-head.
Start from a written verdict
One click — the pair loads with its editorial call on top.
Or compare any two AI models
Pick at least two models to see a scored table, or jump straight to one of the editorial verdicts below.
What this tool finds across every tracked AI model
Run the rules above over every pair of models on this site and they produce 3,257 individual head-to-head calls — one for each benchmark, at each reasoning-effort setting, that both models have a score on. Only 746 of them — 23% — are a confirmed real gap. The remaining 77% resolve to something weaker, and the reasons are the interesting part: 487 land inside the benchmark’s own noise band, 583 have a vendor’s own number on at least one side, 821 compare runs that used different harnesses or different effort settings, and 620 sit on a benchmark this site has graded unfit for ranking at all.
That ratio is worth stating carefully, because a looser version of it flatters this site and is not true: 67% of the 630 model pairs — 424 of them — do have a real gap somewhere, on some benchmark, in some direction. What collapses is the idea that this settles the pair: 221 pairs (35%) separate on one board, tie on another, and return an uncallable result on a third, all at once. “A beats B” is almost never a property of the pair, only of a pair plus a benchmark.
Which benchmark you pick decides more than which models you pick. Among the boards this site still ranks on, HLE returns a real gap 56% of the time, while Terminal-Bench 2.1 manages 6%. Boards this site has stopped ranking on — saturated or contaminated — are excluded from that spread on purpose: they return 0% here by editorial decision, which says nothing about how well they separate models.
Why an AI model comparison comes back undecided
Four different things can stop a comparison short, and they are not interchangeable — one is about statistics, one about who ran the test, one about how it was run, and one about the benchmark itself.
The gap is inside the benchmark’s own noise band
Every benchmark here carries a published threshold derived from its sample size. Below it, a lead is not distinguishable from run-to-run variance, and the tool says Tie — 15% of all calls. On HLE, the largest board this site still ranks on at 2,500 items, anything under 2 points is not a win.
One of the two numbers came from the company selling it
18% of calls have a vendor-reported score on at least one side. That is not an accusation of dishonesty; it is that nobody outside the vendor has reproduced it, so the gap is a claim rather than a result. This check runs before the size of the gap is even measured.
The two scores came from different setups — a different harness, or a different reasoning-effort tier
Two independent evaluators running the same benchmark with different scaffolding — or the same benchmark run at two different reasoning-effort settings — can differ by more than the models do. Where that happens the tool returns Setup-dependent (25% of calls) instead of quietly treating a setup difference as a capability difference.
The benchmark itself is no longer fit to rank on
When a benchmark saturates or its questions leak into training data, this site marks it and stops ranking every pair on it — 19% of all calls, including some with double-digit point spreads. Those are the comparisons most sites still publish a winner for.
How to read a comparison
The Signal column is a pairwise call, not a ranking — it only ever judges the two models you picked against each other, on one benchmark at a time. A real gap on one benchmark says nothing about another; a model can lead on coding and trail on reasoning in the same comparison. Unverified means at least one of the two scores is still vendor-only — treat the gap as a claim, not a confirmed result, until an independent run backs it up. The full rules behind every label are on methodology.