← Current tracker

September 2026 archive

Every AI model release in September 2026

  • 2026-09-04Release

    Anonymous "Omen Alpha" stealth model appears — OpenCode Go subscribers only, no public API

    Announced 2026-09-04 with no lab credit, no model card and no benchmark table. The part that matters is what changed since Ox Alpha: this one has no public API at all — OpenCode's own client on a $10/month Go subscription is the only way to reach it, so no outsider can benchmark it even in principle (Ox Alpha had a public endpoint, which is how one third party ran LiveCodeBench against it). Zero of the eleven boards this site tracks has a row for Omen Alpha, each checked directly on 2026-09-06. OpenCode publishes the price ($0.20 in / $0.66 out / $0.04 cached read per 1M tokens) and a 0-day-retention, no-training-use promise — but not the context window, not the max output, and not the maker. A zhipu/omen-alpha path in OpenCode's own usage data has people calling it Zhipu's second stealth release in nine days; an independent analyst says Xiaomi MiMo-V3-Flash instead. Nobody official has said anything.

    Official source →
  • 2026-09-03Release

    GPT-6 Astra launches — OpenAI's first model to cross its own "Critical" cybersecurity threshold

    Five of the eleven benchmarks this site tracks carry an independent score so far, but only Humanity's Last Exam clears a clean same-tier, same-harness comparison against GPT-5.6 Sol (+5.2 points) — the rest are ties, harness mismatches, or the saturated GPQA Diamond board. Does not replace Sol, Terra, or Luna: OpenAI's own materials never call it a successor, and all three remain sold unchanged.

    Official source →
  • 2026-09-02Release

    Muse Spark 1.3 launches — same price as 1.2, weights decision now explicitly undecided

    Pricing carries over unchanged from Muse Spark 1.2 ($1.25/$4.25 standard, $0.10/$0.20 Contributor tier). Four of the eleven benchmarks this site tracks already carry an independent score one day after launch, all Artificial Analysis at its 'max' tier except LiveBench: GPQA Diamond (93.8%), Humanity's Last Exam (49.1%), Terminal-Bench 2.1 (85.8%), and LiveBench (81.6). Where the comparison against Muse Spark 1.2 is clean, the gain is real: LiveBench gains 3.6 points at a matched reasoning-effort tier, clearing that benchmark's own noise band. HLE's headline +3.6 points mixes tiers the same way GPQA Diamond's does, though — on a matched xhigh-to-xhigh basis the real gain is only about 2.0 points (Terminal-Bench 2.1's +5.6 is a tie; GPQA Diamond is saturated and moot either way). DeepSWE, ARC-AGI-2, LiveCodeBench, SWE-bench Verified, Toolathlon-Verified, Agents' Last Exam, and HMMT Feb 2026 have no Muse Spark 1.3 score yet; Meta's own self-reported 75.4% DeepSWE claim has no independent figure to check it against. Reporting describes Meta as having explicitly not decided whether to open the weights, citing EU AI Act Article 53's systemic-risk carve-out — a more uncertain framing than 1.2's own still-unfulfilled 'open weights coming soon' promise.

    Official source →
  • 2026-09-02Release

    Gemini 3.8 Flash launches — same price as 3.7 Flash, built on its checkpoint, not a new pretraining run

    Google's own model card states plainly "Gemini 3.8 Flash is based on Gemini 3.7 Flash" — pricing is identical at $0.75/$3.75 per MTok (introductory through 2026-12-31), and the knowledge cutoff carries over unchanged at March 2026. A restricted sibling, Gemini 3.8 Flash Cyber, ships the same day for vulnerability detection under Google's invitation-only Fairwind Program. Six of the eleven benchmarks this site tracks already carry an independent score one day after launch — GPQA Diamond (94.44%, vals.ai), LiveCodeBench (89.48%, vals.ai), SWE-bench Verified (80.00%, vals.ai), Terminal-Bench 2.1 (87.6%, Artificial Analysis), Humanity's Last Exam (47.8%, Artificial Analysis), and DeepSWE (74% at high effort, tying Claude Opus 5 for the board's top spot) — none from Google's own claimed numbers.

    Official source →
  • 2026-09-01Release

    Claude Fable 5.1 launches — successor to Fable 5, same price, cheaper cache reads

    Anthropic's own comparison table calls it "Successor to Claude Fable 5" — headline pricing unchanged at $10/$50 per MTok, but cached-read pricing cuts 75% (source: platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1). All six benchmark scores logged here so far are independently sourced, but Terminal-Bench 2.1 alone splits three ways: 91.4% (Artificial Analysis), 85.02%/79.03%-with-fallback-correction (vals.ai's own harness), and no score at all yet from the benchmark's own board, tbench.ai. The refusal-to-fallback mechanism that put substituted answers into some published Fable 5 scores is still present in Fable 5.1, per vals.ai's own disclosure.

    Official source →