AI models we track
Every model that has entered the tracker, from first release to the version that replaced it, with the pricing and context details vendors bury in a launch PDF.
The homepage table only shows the newest model in each line that already has at least one sourced benchmark score, so its rankings stay meaningful. This page is the full catalog — including models with no scores sourced yet and previous versions of lines that have since moved on — so you can see everything that exists, not just what is currently rankable.
License terms and knowledge-cutoff dates live on each model’s own page — they vary too much in length and completeness to fit cleanly into a table row.
All 36 AI models, with pricing and verified score counts
| Context | |||||
|---|---|---|---|---|---|
| Mistral Large 4 Mistral AI | 2026-10-06 | $0.68 / $2.09 | 1M* | 0 independent / 1 total | |
| Gemini 4 Argon Google DeepMind | 2026-09-30 | $4.00 / $20.00introductory rate$2.00 / $10.00 | 1M* | 1 independent / 1 total | |
| GPT-6.1 Sol OpenAI | 2026-09-29 | $2.00 / $10.00 | 1M | 5 independent / 5 total | |
| Claude Sonnet 5.5 Anthropic | 2026-09-28 | $2.00 / $10.00 | 1M | 4 independent / 7 total | |
| MiMo-V2.6-Pro Xiaomi | 2026-09-22 | $0.43 / $0.87 | 1M | 2 independent / 5 total | |
| MiMo-V2.6-Flash Xiaomi | 2026-09-22 | $0.14 / $0.28 | 1M | 1 independent / 4 total | |
| Claude Opus 5.5 Anthropic | 2026-09-22 | $4.00 / $20.00 | 1M | 6 independent / 6 total | |
| GPT-6 Sol OpenAI | 2026-09-22 | $2.00 / $10.00 | 1M | 6 independent / 6 total | Previous version |
| GPT-6 Luna OpenAI | 2026-09-22 | $0.10 / $0.50 | 1M | 5 independent / 6 total | |
| Grok 4.7 xAI (SpaceXAI) | 2026-09-21 | $2.00 / $6.00 | 500K | 3 independent / 4 total | |
| Step 5 Preview StepFun | 2026-09-20 | $1.00 / $2.70 | 1M | 2 independent / 7 total | |
| DeepSeek V4.1 Flash DeepSeek | 2026-09-10 | $0.30 / $1.20off-peak$0.15 / $0.60 | 1M | 3 independent / 7 total | |
| GPT-6 Astra OpenAI | 2026-09-03 | $10.00 / $50.00 | 1M | 7 independent / 7 total | |
| Gemini 3.8 Flash Google DeepMind | 2026-09-02 | $0.75 / $3.75 | 1M | 8 independent / 8 total | |
| Muse Spark 1.3 Meta (Meta Superintelligence Labs) | 2026-09-02 | $1.25 / $4.25Contributor tier$0.10 / $0.20 | 1M | 5 independent / 5 total | |
| Claude Fable 5.1 Anthropic | 2026-09-01 | $10.00 / $50.00 | 1M | 7 independent / 7 total | |
| Tencent Hy4 preview Tencent | 2026-08-28 | $0.83 / $2.50 | 1M | 1 independent / 6 total | |
| GLM-5.3-Flash Zhipu AI (Z.ai) | 2026-08-26 | $0.15 / $0.50 | 1M | 7 independent / 8 total | |
| Qwen3.8-Flash-Next Alibaba | 2026-08-26 | — | 262K | 4 independent / 7 total | |
| GLM-5.3 Zhipu AI (Z.ai) | 2026-08-14 | $1.40 / $4.40 | 1M | 7 independent / 9 total | |
| DeepSeek V4 Pro (0813) DeepSeek | 2026-08-13 | $1.32 / $3.96off-peak$0.66 / $1.98 | 1M | 8 independent / 11 total | |
| Gemini 3.7 Flash Google DeepMind | 2026-08-13 | $0.75 / $3.75 | 1M | 9 independent / 9 total | Previous version |
| Grok 4.6 xAI (SpaceXAI) | 2026-08-12 | $2.00 / $6.00 | 500K | 9 independent / 9 total | Previous version |
| Muse Spark 1.2 Meta (Meta Superintelligence Labs) | 2026-08-05 | $1.25 / $4.25Contributor tier$0.10 / $0.20 | 1M | 7 independent / 7 total | Previous version |
| Qwen3.8-Max Alibaba | 2026-08-03 | $2.00 / $6.00 | 1M | 9 independent / 11 total | |
| DeepSeek V4 Flash (0731) DeepSeek | 2026-07-31 | $0.44 / $1.32off-peak$0.22 / $0.66 | 1M | 9 independent / 9 total | Previous version |
| Claude Opus 5 Anthropic | 2026-07-24 | $5.00 / $25.00 | 1M | 10 independent / 10 total | Previous version |
| Kimi K3 Moonshot AI | 2026-07-16 | $3.00 / $15.00 | 1M | 12 independent / 13 total | |
| GPT-5.6 Sol OpenAI | 2026-07-09 | $4.00 / $20.00 | 1M | 10 independent / 10 total | Previous version |
| GPT-5.6 Luna OpenAI | 2026-07-09 | $0.20 / $1.20Batch API pricing$0.10 / $0.60 | 1M | 8 independent / 8 total | Previous version |
| Claude Sonnet 5 Anthropic | 2026-06-30 | $2.00 / $10.00 | 1M | 9 independent / 9 total | Previous version |
| GLM-5.2 Zhipu AI (Z.ai) | 2026-06-16 | $1.40 / $4.40 | 1M | 11 independent / 12 total | Previous version |
| Claude Fable 5 Anthropic | 2026-06-09 | $10.00 / $50.00 | 1M | 10 independent / 10 total | Previous version |
| MiniMax M3 MiniMax | 2026-06-01 | $0.30 / $1.20 | 1M | 7 independent / 7 total | |
| Claude Opus 4.8 Anthropic | 2026-05-28 | $5.00 / $25.00 | 1M | 12 independent / 13 total | Previous version |
| Gemini 3.1 Pro Preview Google DeepMind | 2026-02-19 | $2.00 / $12.00 | 1M | 11 independent / 12 total |
Second published rate, with the condition attached to it: Muse Spark 1.2 — Contributor tier (opts prompts/completions into Meta's training data); GPT-5.6 Luna — Batch API pricing (async processing, short context); Muse Spark 1.3 — Contributor tier (opts prompts/completions into Meta's training data).
Verified = how many of this model’s sourced benchmark scores were run by an independent third party rather than self-reported by the vendor; “—” means no scores are sourced yet. Prices are US dollars per million tokens, input then output, at the vendor’s headline rate.
How to read this AI model catalog
The verified column is the one worth reading first: it counts how many of a model’s tracked benchmark scores were run by someone other than the company selling it. Two models can sit one row apart with very different answers to that question, and the difference matters more than any single number either of them posts.
Why previous versions stay listed
12 of the 36 models here are previous versions — a newer one has since shipped in the same line, at the same tier. They stay in the catalog because a comparison needs both sides: you cannot ask what a new release actually improved without the model it replaced still being on the record, at its real price and with its real scores. How this site decides what counts as a previous version.
Which labs are covered
14 labs so far: DeepSeek, Anthropic, OpenAI, Google DeepMind, Moonshot AI, Alibaba, Zhipu AI (Z.ai), xAI (SpaceXAI), Meta (Meta Superintelligence Labs), MiniMax, StepFun, Xiaomi, Tencent and Mistral AI. Coverage follows where independent benchmark data exists, not lab size — a model nobody outside its vendor has scored is hard to say anything useful about, whoever built it.
What the catalog adds up to
Across all 36 models, 276 benchmark scores are tracked and 235 of them — 85% — were run by someone other than the vendor. That is not a quality score; it is a measure of how much of this catalog is checkable at all. The 41 scores still resting on a vendor’s own run are marked as such on every page they appear on — each one a place where this site is repeating a claim rather than confirming a result.
Published input pricing runs from $0.10 per million tokens (GPT-6 Luna) to $10.00 (Claude Fable 5, Claude Fable 5.1 and GPT-6 Astra) — a 100× range. Price is the one column here never in dispute: it comes from the vendor’s own page and needs no independent run to confirm, which is more than can be said for anything else in this table. 7 models publish a second rate under conditions the headline price does not mention, and those conditions differ enough per vendor that they live on each model’s own page rather than in a column.
Context windows have converged: 33 of 36 are recorded at a million tokens or more. What that number actually buys differs per vendor, and the full breakdown of which windows are native, extended, or quietly price-tiered is its own piece. One of those is a lab’s figure rather than a vendor’s, marked * below.
* Context window not stated by the vendor. Gemini 4 Argon — not vendor-stated — Artificial Analysis's figure; Google publishes a 1M output limit, not a context window. Mistral Large 4 — vendor-stated (docs.mistral.ai model page: "Context 1M"); Artificial Analysis's comparison summary separately lists 524k — the vendor spec is recorded here, AA's figure noted on the page.
How a model gets listed here
Models ship faster than any one person can write pages for, so this catalog is deliberately not exhaustive. Three kinds get added: a genuinely new release, a lab’s flagship, and a cheap model whose price makes it a real alternative to one. The third category is the one most trackers skip and the one buyers actually weigh, which is why a flash-tier model sits in the same table as a frontier one here.
Nothing is removed once added. A previous version keeps its page, its prices and its scores exactly as they were, because the alternative — deleting the old row when the new one lands — is what makes an upgrade look like an improvement by default. What gets dropped is future editorial effort on it, not the record.