The Model Gap

New AI models — August 2026

14 tracked releases from 6 labs so far, 7 we’d call major. Confirmed-only — no rumors, each row links to an official source.

August 2026

  • 2026-08-14ReleaseMajor

    Z.ai claims +50% coding over GLM-5.2, but only vendor-harness numbers exist and there's no API pricing yet. Verdict pending independent runs.

    Official source →
  • 2026-08-13Release
    Gemini 3.7 Flash reaches GA

    Second Flash GA in a month; intro pricing through 2026-12-31.

    Official source →
  • 2026-08-13ReleaseMajor

    The steepest single-generation jump we've tracked: DeepSWE went from 12.8 (April preview) to 62.7 in four months. The 'beats Opus 4.8' headline is mostly noise on close inspection — see the verdict — but the generational trajectory itself is real.

    Official source →
  • 2026-08-05Deprecation
    Claude Opus 4.1 retired from the API

    Exactly on its announced schedule; recommended replacement is Opus 4.8. Anthropic says it retires models to free capacity for new releases, and has committed to preserving the weights.

    Official source →
  • 2026-08-03ReleaseMajor

    Debuted as the highest-scoring new entry on the independent DeepSWE leaderboard at launch — one of the few claims in this batch that's independently confirmed rather than self-reported.

    Official source →

July 2026

  • 2026-07-24ReleaseMajor

    Same $5/$25 as Opus 4.8 it supersedes. Independent runs back the launch claims: #1 on vals.ai SWE-bench Verified (97.0) — though that benchmark is near-saturated.

    Official source →
  • 2026-07-21Release
    Gemini 3.6 Flash reaches GA

    Workhorse-tier update; Google's Pro line (3.1) is still in preview, so Flash remains its only GA track.

    Official source →
  • 2026-07-16ReleaseMajor

    Debuted at #1 on the independent Frontend Code Arena, ahead of GPT-5.6 Sol and Claude Fable 5 in blind evaluation — a genuinely independent result, not a vendor claim. Open weights carry a bespoke license, not plain MIT/Apache.

    Official source →
  • 2026-07-09Release

    Long-context tier (above ~922K input tokens) costs roughly 2x the standard rate — worth checking your typical prompt length before assuming the headline $5/$30 price applies.

    Official source →

June 2026

  • 2026-06-30ReleaseMajor

    $2/$10 and 1M context at the family's lowest price — and the intro pricing was later made permanent. Six independent scores on our board; sits well below Opus 5/Fable 5 on coding benchmarks, as priced.

    Official source →
  • 2026-06-16Release

    MIT-licensed open weights — the most permissive license in this batch of releases; several competitors use bespoke terms that require a commercial agreement above certain revenue or usage thresholds.

    Official source →
  • 2026-06-15Deprecation
    Claude Sonnet 4 and Opus 4 retired from the API

    Announced April 14, gone June 15 — Anthropic's ~60-day retirement cadence in action. If you pin old model versions, this is the clock you're on.

    Official source →
  • 2026-06-09ReleaseMajor

    A new price tier ($10/$50), not an Opus successor. Independently #1 on Terminal-Bench 2.1 (83.8) — one of the few launch claims a third party has confirmed.

    Official source →

May 2026

  • 2026-05-28Release

    Same headline price as Opus 4.5 ($5/$25 per 1M tokens) — Anthropic held the line on cost while adding an effort-control dial. No longer Anthropic's most capable model as of June 2026, but still the reference point most comparisons use.

    Official source →