New AI model releases — September 2026
46 tracked releases from 15 labs so far, 35 we’d call major. Confirmed-only — no rumors, each row links to an official source.
AI models released in September 2026 (21)
- 2026-09-30ReleaseMajor
Google announces Gemini 4 Argon — frontier model, not yet generally available, and priced at $2/$10 only for an introductory period
Google announced Gemini 4 Argon on 2026-09-30 (20:00 UTC) and two days later, checked 2026-10-02, it is still not generally available: no API model identifier appears in Google's own model list, no general-release date is given, and the model is rolling out first to trusted cyber defenders under the Fairwind Program. The $2/$10 introductory rate doubles to $4/$20 after the intro period, and no vendor-stated context window, knowledge cutoff or reasoning-effort control is published. Two labs have measured it. Artificial Analysis puts it at 57.1 on Humanity's Last Exam without tools, 9.3 points above Gemini 3.8 Flash at that same tier against a 2-point band, though Argon is the only independent HLE row on this site not recorded at max effort, so its column position is not a like-for-like ranking; vals.ai ranks it fifth on Terminal-Bench 4.0 at 57.58% and first on its own Vals Index at 68.90%. Its page is live; what is still unstated by Google is recorded there and below.
Official source →What we can and can’t confirm
Availability, checked 2026-10-02. This is an announcement, not a launch. Google's post is published and says the model is "rolling out soon": access is limited to trusted cyber defenders through the Fairwind Program, and Google says it is participating in the U.S. government's voluntary process for pre-release model access. General release is promised "as soon as possible… starting with paid API customers and Google AI Ultra subscribers" — no date. Google's own API model list, its pricing page, its Vertex AI model list, its changelog and cloud.google.com/release-notes were each checked on 2026-10-01 and none carries an Argon identifier or a Gemini 4 row; the newest model any of them lists is Gemini 3.8 Flash.
Price. Google's post gives $2 per million input tokens and $10 per million output as introductory rates, with cached input at 95% off the input rate — Google does not print the resulting cached figure, so the $0.10 that Artificial Analysis carries is arithmetic, not a Google statement. After the introductory period the rates are $4 input and $20 output. No batch rate, no tier and no long-context surcharge is published. For anyone budgeting against the headline $2/$10, that doubling is the fact to price in.
What Google claims, and what is independently measured. Google's own numbers are coding ones it is not releasing to the public: DeepSWE v1.1 at 77.9% described as new state of the art, AutomationBench (Zapier) first at 51.3%, LVBench state of the art at 91.7%, and CWE-bench v1 tied for first at 68%. Four further wins are named without any figure — the Vals Index, Vals Finance Agent v2, the Harvey Legal Agent Benchmark and Gray Swan IPI. Google's post also states that defenders in the Fairwind Program will get Argon without cyber guardrails, and that Google monitors the model's chain of thought for misalignment.
Two independent labs have measured it, on different benchmarks. Artificial Analysis's model page carries a Humanity's Last Exam score of 57.09% without tools, its Intelligence Index at 53, and a Terminal-Bench 4.0 score of 57.07; the numbers were read from that page's own embedded record, where Argon appears once, at effort tier high, with no other tier published. That high-tier row is the only independent HLE row on this site not recorded at max effort — every other one comes from a max-effort run — so Argon's third-place position in that column describes the column, not a matched ranking. vals.ai's Terminal-Bench 4.0 board lists it fifth at 57.58% (±2.31) on Mini-SWE-agent, and its Vals Index — one of the four benchmarks Google claims to lead without publishing a figure — does show Argon first, at 68.90% (±0.97). Of the seven external boards this site follows, only Artificial Analysis's Humanity's Last Exam board carries Argon — the one score on its model page — and vals.ai, which the site does not track, carries it on two more. LiveBench, DeepSWE, ARC-AGI, Toolathlon, Agents' Last Exam and MathArena's HMMT board were each rendered and read on 2026-10-02 with their model lists populated, and none lists it; five of those six boards have added models released in September, so their silence is not simply staleness. Google's DeepSWE claim is the sharpest case: it reports 77.9% as new state of the art, and the official DeepSWE board, 28 models and updated 2026-09-22, does not list Argon at all.
What the model page does and does not claim. The page exists because Google's announcement is confirmed, and it carries the one tracked score — Humanity's Last Exam at 57.09 from Artificial Analysis. Three things stay unresolved and are stated on the page rather than smoothed over. There is no API model identifier, so a reader still has nowhere to buy it. The context window is not vendor-stated: Google gives a 1M output ceiling, "up from the previous 64K tokens," which is an output limit, while the 1,000,000 in this site's records is Artificial Analysis's figure; the ceiling implies a window of at least 1M, since a model cannot emit more than it accepts, but Google has not said what the window is, and the row added to this site's context-window piece is marked accordingly. And Terminal-Bench 4.0, where both labs have run Argon, is not a benchmark this site follows — it tracks Terminal-Bench 2.1, whose three independent sources have all stopped accepting new models — so those two figures are recorded here rather than stored as scores.
No supersession. Google's post contains no supersession language and never mentions Gemini 3.1 Pro or any other 3.x model; the only predecessor it names is 3.8 Flash Cyber, which Argon "builds on". Nothing tracked here is marked superseded as a result.
Precedent. Anonymous and pre-release listings on this site are recorded as tracker entries without model pages, and the same holds for a flagship: the entry carries the announcement, the disclosure and the reasons, and a page follows once there is something a reader can act on.
- Google — "Gemini 4 Argon: our next era of frontier intelligence" (Koray Kavukcuoglu; published 2026-09-30 20:00 UTC; introductory $2/$10 then $4/$20, 1M output limit, self-reported benchmark claims, no GA date)
- Artificial Analysis — Gemini 4 Argon (High) model page: Intelligence Index 53, HLE 57.09, released 2026-09-30; read from the page's own embedded record, effort tier "high" (checked 2026-10-02)
- vals.ai — Terminal-Bench 4.0 leaderboard: Gemini 4 Argon #5 at 57.58% (±2.31) on Mini-SWE-agent, board updated 2026-09-29 (checked 2026-10-02)
- vals.ai — Vals Index: Gemini 4 Argon first at 68.90% (±0.97), board updated 2026-09-30 (checked 2026-10-02)
- 2026-09-29ReleaseMajor
OpenAI ships GPT-6.1 Sol — same $2/$10 as GPT-6 Sol, one real independent gain, ties Astra on HLE
$2/$10 per million tokens unchanged; cached input halves to $0.10. OpenAI pitches it as near-Astra at a fifth of Astra's price. 3 of 12 tracked benchmarks carry an independent score one day after launch. Against GPT-6 Sol at max effort, Artificial Analysis's Humanity's Last Exam shows a real gain (52.9% vs 47.9%, band 2); LiveBench (81.6 vs 79.3) and AA's Codex-agent DeepSWE run (69.6% vs 69.0%) are ties. Against GPT-6 Astra at max effort, AA's HLE (54.7%), LiveBench (82.2) and AA's DeepSWE run (67.6%) all read as ties. GPT-6 Sol's model page points to "the newer Sol model" but carries no deprecation notice.
Official source → - 2026-09-28ReleaseMajor
Claude Sonnet 5.5 launches — Sonnet 5's $2/$10 unchanged, "not at the capability frontier" by Anthropic's own card
Priced $2/$10 per million tokens, unchanged from Sonnet 5 including cache and batch rates, flat across the 1M window; Anthropic claims 30%+ faster output and up to 30% lower cost per task than Sonnet 5 — its own tests — and its system card says the model "is not at the capability frontier." One day post-launch, 3 of the 12 benchmarks this site tracks carry an independent score. Two same-tier comparisons with Sonnet 5: Humanity's Last Exam is a real gap, 55.0 vs 41.3 (Artificial Analysis, max effort); LiveBench at xHigh is a tie, 77.8 vs 76.0, with Agentic Coding well below Sonnet 5's (39.3 vs 59.4). Sonnet 5's page now carries a "Legacy" badge while the deprecations page lists it "Active" — the same disagreement as Opus 5 — and Sonnet 5 is the visible fallback for Sonnet 5.5's cyber and frontier-LLM safety blocks, so a published 5.5 score can contain Sonnet 5 answers.
Official source → - 2026-09-25ReleaseMajor
Meituan releases LongCat-2.5-Preview — 1.6T parameters, 1M context, no benchmarks published
Meituan's announcement lists 1.6T parameters (~48B active), a 1M-token context, native multimodal input and a focus on long-running agents. The API price, $0.30 input / $1.20 output per million tokens, is labelled a limited-time discount. Meituan published no benchmark numbers, no independent board lists the model as of 2026-09-27, and no weights have been released or announced as of 2026-09-27, unlike LongCat 2.0. With nothing measurable yet, no model page is added.
Official source →What we can and can’t confirm
Announced on Meituan's LongCat X account at 14:16 UTC on 2026-09-25 ("LongCat-2.5-Preview is now live"); the API changelog entry of the same date adds image understanding, coding improvements and compatibility with Claude Code and other agent tools. The model ID is LongCat-2.5-Preview, with a 1M-token context and a 128K maximum output.
Pricing is ¥2 / ¥8 per million input / output tokens ($0.30 / $1.20), with cached input at ¥0.04 ($0.006), shown as a limited-time discounted price with no length tier. One inconsistency in Meituan's own docs: the Retrieve Model API example still shows a text-only modality and a March creation date, which looks stale against the multimodal launch.
Hugging Face's meituan-longcat organisation holds LongCat 2.0 and its quantised versions, but no 2.5 repository as of 2026-09-27. Artificial Analysis lists LongCat 2.0, not 2.5; LiveBench, ARC Prize, the DeepSWE board, tbench.ai, Toolathlon and Agents' Last Exam do not list 2.5 either.
- Meituan LongCat on X — launch post (2026-09-25 14:16 UTC)
- LongCat API platform — changelog entry for 2.5 Preview (checked 2026-09-27)
- LongCat API platform — 2.5 Preview pricing, labelled limited-time discount (checked 2026-09-27)
- Hugging Face — meituan-longcat organisation: LongCat 2.0 repositories, no 2.5 (checked 2026-09-27)
- 2026-09-23ReleaseMajor
Anonymous "Space Bunny Alpha" appears on OpenRouter — free, 1M-token context, maker undisclosed
OpenRouter lists Space Bunny Alpha from September 23 (14:48 UTC, per its public models API) with a 1M-token context window, up to 524,288 output tokens, text, image and video input, and $0 per million tokens through its public API. The listing credits only an anonymous provider, marks reasoning as mandatory across five effort tiers (default max), and carries an expiration date of 2026-10-05. OpenCode Go also lists a "Space Bunny Free" option, free for a limited time. As of 2026-10-01, no standard leaderboard carries it: LiveBench, ARC-AGI-2, DeepSWE, the official Terminal-Bench 2.1 board and Artificial Analysis (models list and changelog through 30 September) all show no row. Two outside sources do publish figures, and their provenance is not the same. AI Benchy, an independent public leaderboard, scores it 7.5 at #118 on a run tested 2026-09-30 — that number is taken from AI Benchy itself. An unaffiliated third-party site publishes its own runs of GPQA Diamond (82.0%, 60-question subset), MMLU-Pro (75%, sample not disclosed) and HLE (46.1%, 300-question subset, 95% CI 40.4-51.8, seven unscored questions), plus a token-efficiency claim. Each is reproduced below with its caveats; none is adopted as a score. No model page is added.
Official source →What we can and can’t confirm
Access terms, checked 2026-09-27 and re-checked 2026-10-01. OpenRouter's public models API gives the listing a creation time of 2026-09-23 14:48 UTC, a 1,000,000-token context window, a 524,288-token output ceiling, text, image and video input with text output, and $0 input and output pricing. Two fields not recorded when this entry was first written: the listing marks reasoning as mandatory with five effort tiers (max, xhigh, high, medium, low; default max), and it carries an expiration date of 2026-10-05 — the free preview is scheduled to end. OpenCode Go separately lists "Space Bunny Free" as free for a limited time; the two listings share a name, but neither page states they serve the same checkpoint.
Leaderboards, checked 2026-10-01. None of the boards this site follows has a row: LiveBench (2026-06-25 release), ARC-AGI-2, DeepSWE (v1.1, updated September 22), the official Terminal-Bench 2.1 board, or Artificial Analysis — neither its models list nor its changelog, which runs through 30 September. The 2026-09-27 version of this entry named the first three of those; the Terminal-Bench 2.1 board and Artificial Analysis's model list are the additions.
AI Benchy, checked 2026-10-01. One separately operated public leaderboard does have it: AI Benchy scores Space Bunny Alpha (high) 7.5 overall, ranked #118, with 13 of 23 tests correct, a 63.8% attempt pass rate, four flaky tests, consistency 8.5, reliability 10.0, zero cost and an average response of 29.38 seconds, on a run tested 2026-09-30. AI Benchy publishes its own methodology page and is not the vendor or the operator of this model; it is a single-maintainer site that carries advertising, so it is independent in the sense that matters here — not affiliated with whoever is serving the model — and not a standards body. Its methodology states that test items and grading internals are kept private to protect integrity, that each model is run repeatedly for stability, and that models are evaluated across reasoning modes. The score is therefore a composite rather than a standardized accuracy, and no third party can check it. By the site's own summary the model ranks first of all in Domain specific and seventeenth of all on Anti-AI Tricks. Two caveats for anyone comparing the headline number. The tiers are not matched: Space Bunny is listed at high, while GPT-6 Sol's high run on the same board scores 9.9 and GPT-5.5's low run scores 9.5, both well above it. And the scale is not a benchmark this site follows. One further caveat applies to every score in this entry: the OpenRouter listing expires on 2026-10-05, so after that date none of these runs can be reproduced, whichever source they came from.
Third-party figures, checked 2026-10-01. A separate site — spacebunnyalpha.com, which states it is not affiliated with OpenRouter, OpenCode or the model's developer — publishes its own runs: GPQA Diamond 82.0%, MMLU-Pro 75%, HLE 46.1%. It also publishes a token-efficiency claim, the one figure here that favours the model: 305,989 output tokens against Qwen3.8 Flash's 913,989 on what it calls the same benchmarks, described as about a third. This site does not adopt any of these as scores, for four stated reasons. First, sample size: GPQA Diamond is a 60-question subset by that site's own account, where 100/√60 is about 12.9 points, so 82.0% spans roughly 69-95 and separates almost nothing; the MMLU-Pro sample is not disclosed; HLE is a 300-question subset with a 95% confidence interval of 40.4-51.8, a standard error of 2.9 points and seven unscored questions, an interval that covers all three GPT reference scores the site lists beside it. Second, the comparison sets are mixed-provenance: Space Bunny's numbers are that site's runs, while the reference scores come from other parties — and on GPQA the mismatch is explicit, since GPT-6 Sol's 95.45% comes from a 198-question audit and GPT-5.5's 93.6% is OpenAI-published, against 82.0% from 60 questions. Third, token counts are run-specific and not comparable across different task sets — AI Benchy's own Space Bunny run, a different suite, logs 181,271 output tokens, so the two token counts do not contradict each other but neither does one support the other. Fourth, GPQA Diamond is already graded D/saturated on this site.
One correction worth recording, because it is the kind of error this site exists to catch. The first version of this entry took Space Bunny Alpha's AI Benchy figures from spacebunnyalpha.com, which reported 7.0 overall, 12 of 22 tests and 62.1% attempt pass rate. AI Benchy's own page reports 7.5, 13 of 23 and 63.8% for a run tested 2026-09-30, and the third-party site places GPT-6 Sol at 6.0 where AI Benchy's own board has its high run at 9.9 — a gap too large to be rounding. Whether that copy reflects an earlier run or a different configuration cannot be determined from what each site publishes. Its AI Benchy figures are not used here; where a primary source exists, this entry cites the primary source. A second omission in that first version is worth naming for the same reason: it left out the token-efficiency claim, which is the one figure in this entry that favours the model.
Identity, checked 2026-10-01. Still no claim from any lab, and neither OpenRouter nor OpenCode names one. AI Benchy's model page independently notes that the official listing requires reasoning and offers low, medium, high, xhigh and max efforts, which matches what OpenRouter reports. The third-party site assesses it as a model fusion router whose leading core-model candidate is GPT-6.1 Sol, with Astra as an alternative, on the basis of a probe that returned Harmony-style channels (analysis, commentary, final, summary) and an OpenAI-style runtime signature; the site records its own status as unconfirmed and says the operator remains unknown. That is one party's method-based assessment, not a reveal. If the router reading is right it adds a caveat to every number above: the scores would describe a routing mixture rather than one model. A tokenizer clue sometimes cited for this listing could not be checked — OpenRouter's tokenizer field is null for every other model tested and reads only "Other" here, so it cannot support or rule out any attribution.
Precedent. Earlier anonymous listings on the same page were later revealed: Ox Alpha as Z.ai's GLM-5.3-Flash, Union Alpha as Unbiased's Pareto 26.9. A reveal would be recorded as a separate entry, as those were, and a tracked model page would follow only if the revealed model meets this site's inclusion criteria.
- OpenRouter — Space Bunny Alpha listing: anonymous provider, 1M context, free (checked 2026-09-27)
- OpenRouter — public models API: created 2026-09-23 14:48 UTC, 524,288 max output, text/image/video input (checked 2026-09-27)
- OpenCode — Go model list: "Space Bunny Free", free for a limited time (checked 2026-09-27)
- spacebunnyalpha.com — unaffiliated third-party site: its own GPQA Diamond, MMLU-Pro and HLE runs, a token-efficiency claim, and a 30 September identity assessment (checked 2026-10-01; the site states it is not affiliated with OpenRouter, OpenCode or the model's developer)
- AI Benchy — independent public leaderboard: Space Bunny Alpha (high) scores 7.5, ranked #118, 13 of 23 tests, 63.8% attempt pass rate, run tested 2026-09-30 (checked 2026-10-01)
- AI Benchy — published methodology: private test items, repeated runs per model, and evaluation across reasoning modes (checked 2026-10-01)
- 2026-09-22ReleaseMajor
GPT-6 Luna launches alongside Sol — cheapest GPT-6 at $0.10/$0.50, ties GPT-5.6 Luna on every clean comparison
Same tiered pricing rule as Sol at one-twentieth Sol's price — and half GPT-5.6 Luna's input rate, 42% of its output rate ($0.20/$1.20, a rate OpenAI itself calls promotional). 3 of 11 tracked benchmarks carry an independent score, and all three tie GPT-5.6 Luna inside the noise band, numerically slightly lower each time: Humanity's Last Exam 38.5% vs 39.5%, LiveBench 72.0 vs 73.6, ARC-AGI-2 59.31% vs 59.6%, same evaluators, max effort. GPT-5.6 Luna is not deprecated; OpenAI's deprecations page lists it as the recommended replacement for two other retiring models.
Official source → - 2026-09-22ReleaseMajor
OpenAI launches GPT-6 Sol — half GPT-5.6 Sol's promotional price, no measurable gain over it yet
$2/$10 per million tokens up to 272K prompt tokens (2x input/cache, 1.5x output above it). OpenAI's table reads "50% cheaper" than GPT-5.6 Sol, but calls GPT-5.6 Sol's $4/$20 rate itself "promotional." 3 of 11 tracked benchmarks carry an independent score, and on the two clean matched comparisons GPT-6 Sol ties its predecessor, numerically slightly lower: Humanity's Last Exam 47.9% vs 49.49% and LiveBench 79.3 vs 81.0, same evaluators, max effort, both inside the noise band. AA's own Codex-agent DeepSWE run (69.0%, close to OpenAI's self-reported 68.8%) uses a different harness from GPT-5.6 Sol's board score. GPT-5.6 Sol carries no deprecation notice.
Official source → - 2026-09-22ReleaseMajor
Claude Opus 5.5 launches — Anthropic's "new leading model" at 20% below Opus 5's per-token price
Priced $4/$20 per million tokens, 20% below Opus 5's $5/$25 per token, flat across the full 1M-token window; Anthropic separately claims 40% lower cost than Opus 5 on typical workloads, its own test. Anthropic calls it "the new leading model," performing "at the level of Claude Fable 5.1 on most work" at 40% of Fable 5.1's price. One day post-launch, 2 of the 11 benchmarks this site tracks carry an independent score, and both back that up: Humanity's Last Exam (61.4%, Artificial Analysis, max effort — the highest AA has recorded, 2.3 points past Fable 5.1, just outside the noise band) and LiveBench (83.2, tied with Fable 5.1's 83.4). Opus 5's own page now carries a "Legacy" badge, but Anthropic's model-deprecations page still lists it "Active" — two official sources disagree.
Official source → - 2026-09-22ReleaseMajor
Xiaomi launches MiMo-V2.6-Flash — 309B open-weight at $0.14/$0.28 per 1M tokens
MIT-licensed 309B / 15B sibling at flat $0.14/$0.28. Zero independent measurements at launch: no Artificial Analysis page (404 checked 2026-09-22), no listing on the seven boards checked live, all four tracked scores Xiaomi's own; same DeepSWE source split as Pro (67.9 on the card vs 65.7 as this RL run's after-training endpoint in the announcement).
Official source → - 2026-09-22ReleaseMajor
Xiaomi launches MiMo-V2.6-Pro — MIT-licensed 1.02T open-weight at $0.435/$0.87 per 1M tokens
Open-weight flagship under MIT: 1.02T total / 42B activated, 1M context, flat $0.435/$0.87. Artificial Analysis independently indexes it at 46, and its Humanity's Last Exam chart lists the model at 49.35 — the one independent score this site tracks for it (corrected 2026-09-23; this entry first said AA published no per-benchmark scores). The seven boards checked live on launch day (tbench.ai, deepswe.datacurve.ai, toolathlon.xyz, snorkel.ai, livebench.ai, arcprize.org, matharena.ai) do not list it, so the four coding and agent numbers are Xiaomi's own. Xiaomi's two official sources also disagree on DeepSWE v1.1 — 71.9 on the model card vs 72.6 as this RL run's after-training endpoint in the announcement's own training write-up; recorded unresolved.
Official source → - 2026-09-21ReleaseMajor
Grok 4.7 launches — same price as Grok 4.6, thinner independent record so far
Priced and speed-matched to Grok 4.6, per xAI's own launch post. Two days post-launch, only 2 of the 8 benchmarks where Grok 4.6 has an independent score carry one for Grok 4.7 — Humanity's Last Exam (43.1%, Artificial Analysis, xhigh) and LiveBench (77.4 overall, xHigh, on the same release Grok 4.6 sits on) — and both read close to Grok 4.6 (42.9% and 78.0 respectively), though the HLE comparison isn't same-tier (Grok 4.6's tracked score used high effort, not xhigh). Grok 4.6 itself did not reach its full 8-benchmark independent record until 5-12 days post-launch, so this is not a like-for-like launch-day comparison either way. xAI's own model card claims 71.0% on DeepSWE, but deepswe.datacurve.ai still shows no Grok 4.7 row, so that figure is unverified; GPQA Diamond, SWE-bench Verified, LiveCodeBench, ARC-AGI-2 and Terminal-Bench 2.1 have no score at all yet. xAI does not call this a replacement for Grok 4.6, which stays listed unchanged on the same pricing page. Parameter count is undisclosed by xAI; a widely circulated third-party 2.1T figure traces to no xAI source and is not used here.
Official source → - 2026-09-20ReleaseMajor
StepFun launches Step 5 Preview — $1.00/$2.70 per 1M tokens, one independent score, open weights promised for October 15
StepFun pitches frontier-level software-engineering performance at a flat $1.00/$2.70 per million tokens. Only one of the six scores this site tracks for it is independent — Humanity's Last Exam without tools, 46.48, run by Artificial Analysis — while the coding and agent numbers (DeepSWE 67.7, Terminal-Bench 2.1 85.0, Toolathlon-Verified 74.1) are StepFun's own and none of the boards behind those benchmarks lists the model yet. The open weights are a plan for October 15, not a release: the model is hosted-only, and the third-party copies posted on Hugging Face are not a StepFun release.
Official source →What we can and can’t confirm
Timeline. Artificial Analysis had a page for Step 5 Preview, with a 2026-09-18 release label, in an archive capture from 05:44 UTC on September 19 — at least 21 hours before StepFun's announcement post on X at 03:15 UTC on September 20 (timestamp derived from the post's ID). StepFun's own pages carry no publication date. After launch StepFun edited two things without a changelog: the docs' maximum output went from 1M tokens to 64K tokens, and the Artificial Analysis-sourced rows of its launch table (GDPval-AA, AA-Briefcase) were re-issued under new benchmark versions with new numbers. The rows this site tracks did not change across the three versions of the page that could be compared.
Weights. StepFun's statement is "The model will be released with open weights on October 15" — no year, license or repository. Its official Hugging Face repository (stepfun-ai/Step-5-Preview-BF16) was not publicly visible to anonymous visitors on September 21 and 22 (Hugging Face returned a 404 page and a 401 from its public API). Since September 20, third parties have posted copies labeled Step-5-Preview-BF16; the first, at 05:07 UTC, is marked as a duplicate of StepFun's repository, and at least one carries an uploader-written README that presents itself as StepFun's and contradicts StepFun's own documentation. StepFun's site, docs and X account say nothing about them (checked 2026-09-22). This site treats the model as proprietary until StepFun publishes weights and a license.
Independent coverage. Artificial Analysis is the only tracked evaluator that lists the model, and only on Humanity's Last Exam among this site's eleven benchmarks (its Terminal-Bench score is on version 4.0, which this site does not track). As of September 21 the DeepSWE, Terminal-Bench 2.1, Toolathlon, Agents' Last Exam, ARC-AGI-2, LiveBench and MathArena boards did not list it, and vals.ai's GPQA Diamond, SWE-bench Verified and LiveCodeBench pages no longer run new models.
- StepFun — Step 5 Preview announcement and launch table (read 2026-09-22)
- StepFun — API pricing, international USD list (checked 2026-09-21)
- StepFun — Step 5 Preview model page: context window, maximum output, effort levels (checked 2026-09-21)
- StepFun on X — launch post (dated 2026-09-20 03:15 UTC by its ID)
- Artificial Analysis — Step 5 Preview model page (checked 2026-09-21)
- Artificial Analysis — Humanity's Last Exam leaderboard (checked 2026-09-22)
- Hugging Face — stepfun-ai/Step-5-Preview-BF16, not publicly visible (checked 2026-09-22)
- Wayback Machine — Artificial Analysis Step 5 Preview page, earliest capture (2026-09-19 05:44 UTC)
- Wayback Machine — StepFun model docs at launch, maximum output 1M tokens (2026-09-20 03:57 UTC)
- Wayback Machine — StepFun model docs after the edit, maximum output 64K tokens (2026-09-21 13:44 UTC)
- StepFun — China-region price list in yuan (checked 2026-09-21)
- Terminal-Bench 2.1 leaderboard on tbench.ai — no Step 5 Preview entry (checked 2026-09-21)
- DeepSWE leaderboard, v1.1 — no Step 5 Preview entry (checked 2026-09-21)
- Toolathlon-Verified leaderboard — no Step 5 Preview entry (checked 2026-09-21)
- Agents' Last Exam leaderboard (Snorkel AI) — no Step 5 Preview entry (checked 2026-09-21)
- ARC Prize leaderboard — no Step 5 Preview entry (checked 2026-09-21)
- LiveBench — no Step 5 Preview entry (checked 2026-09-21)
- MathArena — no Step 5 Preview entry (checked 2026-09-21)
- 2026-09-17Release
Union Alpha revealed as Unbiased's Pareto 26.9 — a multi-model blend, now paid at $2.50/$7.50 per 1M tokens
OpenRouter's own stealth page now says the listing was "revealed to be Pareto by Unbiased" and that the free stealth period has ended. Unbiased describes Pareto as a blend that calls multiple LLMs in parallel on every request and synthesizes the results — explicitly not a router — but names none of the models inside it, so this site cannot attribute a score to a single lab, weights file or checkpoint. Its model card publishes five numbers (DeepSWE 74, Terminal-Bench 4.0 51, MMMU-Pro 78, HLE without tools 49, ArXivMath 88) that are Unbiased's own; its how-it-works page points to a public evaluation harness, but none of the numbers has been run independently, and none is added here.
Official source →What we can and can’t confirm
OpenRouter's stealth page, checked 2026-09-21, still lists Union Alpha (added September 16, 262K context window) and now carries the note that it was "revealed to be Pareto by Unbiased" and that "the free stealth period has ended; use Pareto" at openrouter.ai/unbiased/pareto. OpenRouter's own model listing for Pareto (unbiased/pareto) carries a creation time of 23:02 UTC on September 17, a 262,144-token context window, 131,072 maximum output tokens and the same prices. The reveal came on X on the evening of September 17 UTC: the @unionalphaai account's post is timestamped 23:03 UTC, Unbiased's own account posted at 23:05 UTC and OpenRouter's announcement is timestamped 23:24 UTC (all times derived from the posts' IDs).
Unbiased's model card, checked 2026-09-21, lists Pareto 26.9 as "our blended AI model" with the identifier pareto, text and image input, and a price of $2.50 per million input tokens, $0.25 for cached input and $7.50 for output. Its how-it-works page says "multiple LLMs are engaged in parallel on every request, and their results are synthesized dynamically based on the task," that the blend "includes frontier and open models," and that Pareto "is not a 'model router.'" It does not say which models, and points to a public evaluation harness (github.com/circuitandchisel/pareto-evals). The card states that measured task costs and a composite score "have not been published for this release," and the how-it-works page says latency comparisons have not been published either.
LiveBench, an independent board this site tracks, lists "Union Alpha" under the organization "Stealth" with an Overall score of 76.1 in data files last modified on September 17; it has not relabeled the row, and OpenCode's Go documentation no longer lists the model (checked 2026-09-21). A LiveBench row for a blend measures the blend, not any one lab's model.
This stays a tracker event rather than a model page. Unbiased describes Pareto as a system of several open and frontier models and does not name them, so there is no single lab, license or checkpoint this site can attach a benchmark to, and the five numbers on its card are Unbiased's own, unreplicated by anyone independent. Two of them share a name with benchmarks this site tracks (DeepSWE, and HLE without tools), but sharing a name is not the same as running under the same harness, so they are not converted into score rows. It is rated minor: the free listing became a paid endpoint, but there is no independent measurement of it and it adds no model to the tracker.
- OpenRouter — stealth listing page: Union Alpha added September 16, 262K context, reveal note and end of the free stealth period (checked 2026-09-21)
- Unbiased — Pareto 26.9 model card: identifier, modalities, pricing and the five published scores (checked 2026-09-21)
- Unbiased — how a Pareto answer gets made: parallel multi-model design, "not a model router" (checked 2026-09-21)
- OpenRouter — reveal announcement on X (post dated 2026-09-17 23:24 UTC by its ID)
- Union Alpha account (@unionalphaai) on X — reveal post (dated 2026-09-17 23:03 UTC by its ID)
- Unbiased on X — post about the reveal (dated 2026-09-17 23:05 UTC by its ID)
- OpenRouter — Pareto listing (created 2026-09-17 23:02 UTC per its public API)
- Unbiased — pareto-evals, the public evaluation harness its how-it-works page cites
- LiveBench — Union Alpha row under organization "Stealth" (data files modified 2026-09-17)
- OpenCode — Go documentation, no longer lists Union Alpha (checked 2026-09-21)
- 2026-09-16ReleaseMajor
Anonymous "Union Alpha" appears on OpenRouter — free public API, maker undisclosed
Update, 2026-09-17: revealed as Unbiased’s Pareto 26.9 — see the September 17 entry. What follows is the launch-day record. OpenRouter lists Union Alpha from September 16 with a 262K context window and $0 per million input and output tokens, available through its public API. The listing describes a multimodal model for research, coding and agentic work, but credits only an anonymous third-party provider. OpenCode Go also lists it as free for a limited time. Access is verifiable; model identity and benchmark performance are separate questions. No model page or scores are added on the strength of this listing.
Official source →What we can and can’t confirm
OpenRouter's stealth page, checked 2026-09-17, dates Union Alpha to September 16 and offers it through the OpenRouter API as stealth/union-alpha. Its displayed context capacity is 262K and both input and output are listed at $0 per million tokens. Those are the listing's access terms, not an independently measured capability result or a guarantee that the preview will stay free. OpenRouter identifies the developer and operator only as an anonymous third-party provider; OpenRouter itself is the router, not the model's maker.
OpenCode's Go documentation, checked the same day, lists Union Alpha Free with model ID union-alpha, a limited-time free offer, no training use and zero-day retention. Those data-handling terms belong to the OpenCode offering; they should not be assumed to describe requests sent through OpenRouter.
This is a tracker event rather than a model page: a public endpoint is enough to document availability, but not enough to fill in a maker, license or benchmark result. The event is marked major for the access change in the anonymous-model story, not for an asserted performance lead. The sources checked for this entry do not establish a relationship to Omen Alpha, which is no longer listed in the current Go documentation. A catalog change alone cannot establish a rename or justify carrying scores from one name to the other.
- 2026-09-10ReleaseMajor
DeepSeek V4.1 Flash launches — replaces V4 Flash; V4 Pro API remains available
Status corrected 2026-09-17 against DeepSeek's September 10 changelog: V4 Flash and V4 Flash Vision Exp are retired, with their API names temporarily routed to V4.1 Flash. V4 Pro service continues after September 14 with billing unchanged; the prior claim here that all Pro requests would redirect on that date is withdrawn. V4 Flash remains superseded on this site; V4 Pro is current. Pricing drops to $0.15/$0.60 off-peak (from V4 Flash's $0.22/$0.66, V4 Pro's $0.66/$1.98), and the architecture changes substantially — a new Causal Encoder-Decoder design DeepSeek says cuts KV-cache size roughly 4x versus V4 Flash, plus native multimodal input. The performance claims lack independent confirmation at the launch check: as of this release, vals.ai, tbench.ai, Artificial Analysis, deepswe.datacurve.ai, toolathlon.xyz, arcprize.org, livebench.ai and matharena.ai all checked live and none list this model yet — every score on its page is DeepSeek's own number.
Official source → - 2026-09-09Price change
GLM-5.3-Flash launch promo ends — back to $0.15/$0.50 per 1M tokens
The 50% launch discount ($0.075 / $0.25) expired on schedule at 24:00 UTC+8, exactly as Z.ai's pricing page had stated from day one; the $0.15 / $0.50 list rate is now the live price, confirmed against OpenRouter's price feed on 2026-09-22. Nothing about the model changed — only the price.
Official source → - 2026-09-04ReleaseMajor
Anonymous "Omen Alpha" appeared on OpenCode Go — no longer listed as of September 17
Omen Alpha was announced on 2026-09-04 as exclusive to OpenCode Go subscribers, without a named maker or model card. Update checked 2026-09-17: OpenCode's Go documentation no longer lists Omen Alpha; it now lists Union Alpha, also available through a public OpenRouter API. Neither listing establishes that Union Alpha is a renamed Omen Alpha. The earlier claim here that outsiders could not benchmark Omen Alpha was too strong: a third-party llmbench run was already documented, although it was not a result from a benchmark tracked by this site.
Official source →What we can and can’t confirm
Historical access record, checked 2026-09-06: OpenCode's Go documentation listed Omen Alpha at $0.20 per 1M input tokens, $0.66 per 1M output and $0.04 per 1M cached reads, with zero-day retention and no training use. These are the terms recorded at that check, not a current offer: Omen Alpha is absent from the Go documentation rechecked on 2026-09-17.
What the September 6 documentation did not say is most of what a model page here would need. No context window. No maximum output. No parameter count, licence, or lab. The 500,000-token context and 128,000-token max output that appear in write-ups about this model were not figures published in that documentation — they come from third-party aggregators such as models.dev and modelcompare.dev, which is a different claim with a different provenance, and this entry does not restate them as spec.
Access and verification are different questions. At launch, Omen Alpha was advertised as exclusive to OpenCode Go subscribers, unlike Ox Alpha's public OpenRouter listing. That made an ordinary public-API benchmark run less straightforward, but did not make outside testing impossible. The llmbench report below is itself a counterexample to this entry's earlier claim that verification was structurally blocked. Restricted access does not, by itself, make every reported result unfalsifiable.
At the 2026-09-06 leaderboard check, none of the eleven benchmarks tracked by this site had an Omen Alpha row. That dated observation is not a fresh September 17 leaderboard sweep. A disclosed-method third-party test can still exist outside those boards, and should be evaluated on its own methodology rather than treated as a tracked benchmark score.
One disclosed-method test does exist, and it is worth reading for what its own author says about it rather than for its headline. A developer publishing as zephel01 ran a self-made SWE-Bench-style suite (llmbench, MIT-licensed on GitHub) and reports 95.0% resolved on a 60-question first stage and 81.2% on a 16-question grandmaster stage. The line that travelled — that Omen Alpha, Ox Alpha and GLM-5.3-Flash all failed the identical question — is real, but the same author states plainly that t103 and t105 are "highly likely ... flaws in the benchmark side" rather than model limitations, and that only one of Omen Alpha's three hardest-stage failures is attributable to the model at all. A shared failure on a question the benchmark's own author suspects is broken is weak evidence of shared lineage, and it is being repeated as strong evidence.
Attribution is two competing guesses and no confirmation. OpenCode's usage-data page files the model under a path reading zhipu/omen-alpha, which if it means what it appears to mean would make this Zhipu's second anonymous release in nine days, following the Ox Alpha → GLM-5.3-Flash reveal on 2026-08-26. Against that, an independent analyst has argued publicly for Xiaomi MiMo-V3-Flash. Neither Zhipu, Xiaomi nor OpenCode has said anything on the record. The Ox Alpha precedent is a caution in both directions: the community's leading guess there did turn out right, but it cycled through Xiaomi and StepFun first, and the reveal settled product identity without settling checkpoint identity — LiveBench's ox-alpha row is still unrenamed as of 2026-09-06, eleven days after Z.ai claimed the model, which is why this site still excludes that row rather than crediting it to GLM-5.3-Flash.
Update checked 2026-09-17: OpenCode Go now lists Union Alpha as free for a limited time and no longer lists Omen Alpha. OpenRouter's stealth page lists Union Alpha with public API access, but no Omen Alpha entry. These are observable catalog changes, not proof of a rename, shared checkpoint, shared maker, or the date Omen Alpha access ended. A confirmed identity announcement or a documented evaluation would justify a further update; this site does not transfer Omen Alpha results to Union Alpha.
- OpenCode — Go documentation; historical Omen terms checked 2026-09-06, Omen absent and Union listed on recheck 2026-09-17
- OpenCode on X — the announcement, 2026-09-04: "Omen Alpha (new stealth model) Exclusively for OpenCode Go subscribers"
- OpenRouter — stealth page lists Union Alpha and Ox Alpha, not Omen Alpha (checked 2026-09-17)
- zephel01 — self-made llmbench run, including the author's own caveat that the shared failure question is likely a benchmark flaw
- modelcompare.dev — third-party aggregator listing the 500K context figure that OpenCode itself does not publish
- Startup Fortune — reporting on the zhipu/omen-alpha usage-data path and the resulting attribution guesses
- OrcaRouter — the competing Xiaomi MiMo-V3-Flash hypothesis
- 2026-09-03ReleaseMajor
GPT-6 Astra launches — OpenAI's first model to cross its own "Critical" cybersecurity threshold
Five of the eleven benchmarks this site tracks carry an independent score so far, but only Humanity's Last Exam clears a clean same-tier, same-harness comparison against GPT-5.6 Sol (+5.2 points) — the rest are ties, harness mismatches, or the saturated GPQA Diamond board. Does not replace Sol, Terra, or Luna: OpenAI's own materials never call it a successor, and all three remain sold unchanged.
Official source → - 2026-09-02ReleaseMajor
Muse Spark 1.3 launches — same price as 1.2, weights decision now explicitly undecided
Pricing carries over unchanged from Muse Spark 1.2 ($1.25/$4.25 standard, $0.10/$0.20 Contributor tier). Four of the eleven benchmarks this site tracks already carry an independent score one day after launch, all Artificial Analysis at its 'max' tier except LiveBench: GPQA Diamond (93.8%), Humanity's Last Exam (49.1%), Terminal-Bench 2.1 (85.8%), and LiveBench (81.6). Where the comparison against Muse Spark 1.2 is clean, the gain is real: LiveBench gains 3.6 points at a matched reasoning-effort tier, clearing that benchmark's own noise band. HLE's headline +3.6 points mixes tiers the same way GPQA Diamond's does, though — on a matched xhigh-to-xhigh basis the real gain is only about 2.0 points (Terminal-Bench 2.1's +5.6 is a tie; GPQA Diamond is saturated and moot either way). DeepSWE, ARC-AGI-2, LiveCodeBench, SWE-bench Verified, Toolathlon-Verified, Agents' Last Exam, and HMMT Feb 2026 have no Muse Spark 1.3 score yet; Meta's own self-reported 75.4% DeepSWE claim has no independent figure to check it against. Reporting describes Meta as having explicitly not decided whether to open the weights, citing EU AI Act Article 53's systemic-risk carve-out — a more uncertain framing than 1.2's own still-unfulfilled 'open weights coming soon' promise. (Update 2026-09-29: this site now tracks twelve benchmarks, and Artificial Analysis has since revised the three max-tier figures to 93.5, 48.7 and 84.3, which trims HLE's headline gain to +3.2 and Terminal-Bench 2.1's to +4.1 — neither verdict changes.)
Official source → - 2026-09-02ReleaseMajor
Gemini 3.8 Flash launches — same price as 3.7 Flash, built on its checkpoint, not a new pretraining run
Google's own model card states plainly "Gemini 3.8 Flash is based on Gemini 3.7 Flash" — pricing is identical at $0.75/$3.75 per MTok (introductory through 2026-12-31), and the knowledge cutoff carries over unchanged at March 2026. A restricted sibling, Gemini 3.8 Flash Cyber, ships the same day for vulnerability detection under Google's invitation-only Fairwind Program. Six of the eleven benchmarks this site tracks already carry an independent score one day after launch — GPQA Diamond (94.44%, vals.ai), LiveCodeBench (89.48%, vals.ai), SWE-bench Verified (80.00%, vals.ai), Terminal-Bench 2.1 (87.6%, Artificial Analysis), Humanity's Last Exam (47.8%, Artificial Analysis), and DeepSWE (74% at high effort, tying Claude Opus 5 for the board's top spot) — none from Google's own claimed numbers.
Official source → - 2026-09-01ReleaseMajor
Claude Fable 5.1 launches — successor to Fable 5, same price, cheaper cache reads
Anthropic's own comparison table calls it "Successor to Claude Fable 5" — headline pricing unchanged at $10/$50 per MTok, but cached-read pricing cuts 75% (source: platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1). All six benchmark scores logged here so far are independently sourced, but Terminal-Bench 2.1 alone splits three ways: 91.4% (Artificial Analysis), 85.02%/79.03%-with-fallback-correction (vals.ai's own harness), and no score at all yet from the benchmark's own board, tbench.ai. The refusal-to-fallback mechanism that put substituted answers into some published Fable 5 scores is still present in Fable 5.1, per vals.ai's own disclosure.
Official source →
AI models released in August 2026 (13)
13 events (10 major): Tencent Hy4 preview — 770B open-weight at $0.834/$2.501 per 1M tokens, added to tracker retroactively, Qwen3.8-Flash-Next ships open-weight — a Qwen4-architecture preview under a Qwen3.8 label, Ox Alpha revealed as GLM-5.3-Flash — GA, MIT license, $0.15/$0.50 per 1M tokens, GPT-5.6 Sol price drops to $4.00/$20.00 per 1M tokens (promotional), Anonymous "Ox Alpha" stealth model appears on OpenRouter, GLM-5.3 API goes live — $1.40/$4.40 per 1M tokens, same price as GLM-5.2, GLM-5.3 released — subscription-only for now, Gemini 3.7 Flash reaches GA, DeepSeek V4 Pro (0813) exits preview, Grok 4.6 released, Muse Spark 1.2 launches — a coding-focused point release in Meta's closed model line, Claude Opus 4.1 retired from the API, Qwen3.8-Max launches — full month →
AI models released in July 2026 (6)
6 events (4 major): DeepSeek V4 Flash 0731 ships with a changelog line, no launch post, Claude Opus 5 released, Gemini 3.6 Flash reaches GA, Kimi K3 launches (API); open weights follow July 26-27, GPT-5.6 Luna launches — cheapest tier of the GPT-5.6 family, GPT-5.6 Sol reaches general availability — full month →
AI models released in June 2026 (5)
5 events (2 major): Claude Sonnet 5 released, GLM-5.2 reaches general release, Claude Sonnet 4 and Opus 4 retired from the API, Claude Fable 5 — first Mythos-class model goes GA, MiniMax M3 — open-weight value play, added to tracker retroactively — full month →
AI model release tracker FAQ
How many AI models were released in September 2026?
21 tracked events in September 2026, out of 46 since this tracker started. Each one is a release, price change or deprecation we could confirm against an official source — announcements we could not confirm are not listed at all.
Which AI model releases here actually matter?
35 of the 46 tracked events are marked major. That label is a judgment, not a vendor claim: it means the release changed something measurable — a price, a capability tier, or a benchmark position — rather than that it was heavily promoted.
Where do these AI model release dates come from?
Official sources only — vendor blogs, API changelogs, model cards. Every row links to one. Where a model shipped with no announcement, the date is reconstructed from version strings and repository timestamps and the row says so, because a date we inferred and a date a vendor published are not the same fact.
What counts as an event on this AI model tracker?
release, price change, deprecation — across 15 labs so far. Deprecations and price changes are tracked alongside launches because for anyone already running a model in production they matter more than a new name does.