August 2026 archive
August is the busiest month tracked here — seven events, four of them major. DeepSeek V4 Pro's 2026-08-13 launch chart claimed a win over Claude Opus 4.8 that mostly turns out to be noise (see the verdict); the real story in its own column is a roughly fifty-point jump on DeepSWE from its prior release. GLM-5.3 released on 2026-08-14 and its API followed on 2026-08-19 at an unchanged price — every benchmark number attached to it was vendor-only at launch, and Artificial Analysis has since run two of its five tracked scores independently.
Every AI model release in August 2026
- 2026-08-26Release
Qwen3.8-Flash-Next ships open-weight — a Qwen4-architecture preview under a Qwen3.8 label
Don't confuse it with the bare 'Qwen3.8-Flash' (no suffix) — Alibaba's own model card describes that as a separate, not-yet-released hosted SKU with production features this preview lacks. Flash-Next is open-weight only, no hosted API price exists yet, and all 4 tracked benchmark scores are Alibaba's own self-reports — none of 8 independent boards checked (Artificial Analysis, vals.ai, ARC Prize, LiveBench, MathArena, DeepSWE, Toolathlon, Snorkel AI) had added it as of launch day.
Official source → - 2026-08-26Release
Ox Alpha revealed as GLM-5.3-Flash — GA, MIT license, $0.15/$0.50 per 1M tokens
Z.ai confirms it outright: 'we tested GLM-5.3-Flash anonymously as ox-alpha ... to gather user feedback.' That settles product identity, not checkpoint identity — LiveBench's 'ox-alpha-max' row (69.2 overall) is still unrenamed as of today, so this site excludes it rather than crediting it to the GA release, the same caution this site applied to a past DeepSeek V4 Pro checkpoint mismatch. Real paid pricing landed the same day as the reveal (promotional $0.075/$0.25 through 2026-09-09, then $0.15/$0.50), a day earlier than this tracker guessed on 2026-08-20.
Official source → - 2026-08-26Price change
GPT-5.6 Sol price drops to $4.00/$20.00 per 1M tokens (promotional)
Down from $5.00/$30.00 checked here on 2026-08-20 — OpenAI's own pricing page labels this a promotional rate guaranteed "at least through November 21, 2026," not a permanent cut. Long-context tier (>272K input) is $8.00/$30.00, still 2x input / 1.5x output of the new standard rate.
Official source → - 2026-08-20Release
Anonymous "Ox Alpha" stealth model appears on OpenRouter
Listed by an anonymous third-party provider — OpenRouter isn't the maker, and no lab has claimed it. 1M-token context, free during a promotional preview week (real pricing expected around 2026-08-27). Zero scores on any benchmark board we track (Artificial Analysis, vals.ai, Arena.ai, Toolathlon, Terminal-Bench) as of 2026-08-22 — rechecked, unchanged. Independent tokenizer fingerprinting now leans toward Zhipu/Z.ai's GLM-5.3 lineage over the earlier Xiaomi MiMo theory, with a third guess (StepFun) surfacing 2026-08-22 — still nobody official has said anything.
Official source → - 2026-08-19Price change
GLM-5.3 API goes live — $1.40/$4.40 per 1M tokens, same price as GLM-5.2
The pricing is the verifiable part: identical to GLM-5.2, cached input $0.26. The launch chart's numbers (88.2 Terminal-Bench 2.1, 66.9 DeepSWE) are still all vendor-run — every score on our GLM-5.3 page stays Unverified until an independent board reruns it.
Official source → - 2026-08-14Release
GLM-5.3 released — subscription-only for now
Z.ai claims +50% coding over GLM-5.2, but only vendor-harness numbers exist and there's no API pricing yet. Verdict pending independent runs.
Official source → - 2026-08-13Release
Gemini 3.7 Flash reaches GA
Second Flash GA in a month; intro pricing through 2026-12-31.
Official source → - 2026-08-13Release
DeepSeek V4 Pro (0813) exits preview
The steepest single-generation jump we've tracked: DeepSWE went from 12.8 (April preview) to 62.7 in four months. The 'beats Opus 4.8' headline is mostly noise on close inspection — see the verdict — but the generational trajectory itself is real.
Official source → - 2026-08-12Release
Grok 4.6 released
Independent boards largely back the launch claims: rank 3 on GPQA Diamond and rank 4 on SWE-bench Verified (vals.ai), and DeepSWE came in above xAI's own number. Agent-board coverage (tbench/Toolathlon/Snorkel) still pending.
Official source → - 2026-08-05Release
Muse Spark 1.2 launches — a coding-focused point release in Meta's closed model line
Independent scores diverge sharply from Meta's own chart: Terminal-Bench 2.1 spans a 13-point range across three sources (Meta 82.9%, Artificial Analysis 80.15%, vals.ai 69.66%), and DeepSWE shows a smaller +4.3pp vendor-favorable gap (Meta 59.3% vs. the independent board's 55%). Open weights promised 'coming soon' on 2026-08-10 remain unreleased as of 2026-08-26 — Meta open-weighted a separate, smaller model (Muse Glimmer) that same day instead.
Official source → - 2026-08-05Deprecation
Claude Opus 4.1 retired from the API
Exactly on its announced schedule; recommended replacement is Opus 4.8. Anthropic says it retires models to free capacity for new releases, and has committed to preserving the weights.
Official source → - 2026-08-03Release
Qwen3.8-Max launches
Debuted as the highest-scoring new entry on the independent DeepSWE leaderboard at launch — one of the few claims in this batch that's independently confirmed rather than self-reported.
Official source →