July 2026 archive
July brought five releases, two of them agentic standouts: Kimi K3 (2026-07-16) debuted #1 on the independent Frontend Code Arena, and DeepSeek V4 Flash 0731 shipped with no announcement post yet gained 19.8 independently-confirmed points on Toolathlon-Verified over its predecessor. Claude Opus 5 arrived on 2026-07-24 at Opus 4.8's own $5/$25 price. GPT-5.6 Sol's headline $5/$30 rate only holds under 272K input tokens — its input price doubles above that threshold.
Every AI model release in July 2026
- 2026-07-31Release
DeepSeek V4 Flash 0731 ships with a changelog line, no launch post
Announced only as a changelog line: +19.8 points on Toolathlon-Verified over the previous Flash, independently confirmed, at the cheapest per-test cost on the SWE-bench board. MIT open weights.
Official source → - 2026-07-24Release
Claude Opus 5 released
Same $5/$25 as Opus 4.8 it supersedes. Independent runs back the launch claims: #1 on vals.ai SWE-bench Verified (97.0) — though that benchmark is near-saturated.
Official source → - 2026-07-21Release
Gemini 3.6 Flash reaches GA
Workhorse-tier update; Google's Pro line (3.1) is still in preview, so Flash remains its only GA track.
Official source → - 2026-07-16Release
Kimi K3 launches (API); open weights follow July 26-27
Debuted at #1 on the independent Frontend Code Arena, ahead of GPT-5.6 Sol and Claude Fable 5 in blind evaluation — a genuinely independent result, not a vendor claim. Open weights carry a bespoke license, not plain MIT/Apache.
Official source → - 2026-07-09Release
GPT-5.6 Luna launches — cheapest tier of the GPT-5.6 family
At $0.20/$1.20 per 1M tokens it undercuts sibling GPT-5.6 Sol by roughly 3-29x per task while trailing by only a few points on shared benchmarks (93.00% vs. 96.20% on SWE-bench Verified). All 8 tracked scores are independently sourced, though Terminal-Bench 2.1 reads 80.9% on Artificial Analysis versus 79.03% on vals.ai's own run of the identical model.
Official source → - 2026-07-09Release
GPT-5.6 Sol reaches general availability
Long-context tier (above ~922K input tokens) costs roughly 2x the standard rate — worth checking your typical prompt length before assuming the headline $5/$30 price applies.
Official source →