Zhipu AI (Z.ai)
GLM-5.3-Flash
Its downloadable safetensors on Hugging Face (MIT license, 1,048,576-token context) settle the license question the site can verify directly, but the score sheet is less settled: Artificial Analysis's own same-day page instead calls it proprietary, text-only, and 400k-context, and of the six benchmark scores gathered so far, only four are independently measured — the other two, Toolathlon Verified and Agents' Last Exam, are Z.ai's own numbers only.
GLM-5.3-Flash’s 6 benchmark scores on this page were verified against their sources on or after 2026-08-26.
- Released
- 2026-08-26
- License
- open-weights
- Context window
- 1M tokens
- Knowledge cutoff
- Not disclosed
The verified record
Against the 88 head-to-head comparisons GLM-5.3-Flash shares with other tracked models: 15 real gaps, 23 inside the noise band, and 50 we will not call.
A gap counts for GLM-5.3-Flash only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — GLM-5.3-Flash trails on 13 of them.
HLE · no toolsReasoning
±2 is noiseDeepSWELong-horizon coding
±9.5 is noiseNo verdict for GLM-5.3-Flash anywhere on Terminal-Bench 2.1, Agents' Last Exam, Toolathlon-Verified (nothing independently confirmed on both sides); GPQA Diamond (saturated).
- None of GLM-5.3-Flash’s agentic comparisons are independently confirmed on both sides yet.
GLM-5.3-Flash API pricing
$0.15 in / $0.50 out per 1M tokens — official pricing
GLM-5.3-Flash is one of 3 Zhipu AI (Z.ai) models tracked on this site, at these official list prices.
| Model | In / 1M | Out / 1M |
|---|---|---|
| GLM-5.3-Flash | $0.15 | $0.50 |
| GLM-5.3 | $1.40 | $4.40 |
| GLM-5.2(superseded) | $1.40 | $4.40 |
GLM-5.3-Flash benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise | |
| GPQA Diamondsaturated[2] Expert science Q&A — not ranked at any gap size | |
| Terminal-Bench 2.1[3] Terminal ops · ±10.6 is noise | |
| DeepSWE[4] Long-horizon coding · ±9.5 is noise | |
| Toolathlon-Verified[5] Multi-tool chores · ±9.7 is noise | |
| Agents' Last Exam[6] Professional work · ±3.2 is noise |
Who ran these numbers: 4 of 6 independent — artificialanalysis.ai (3), deepswe.datacurve.ai (1); vendor self-reported (2).
- HLE: Standard HLE (no tools), independently measured by Artificial Analysis. Exact value from the page's underlying data (0.398517...); AA's on-screen rounded display shows 40%. Distinct from Z.ai's self-reported 'HLE w/ Tools' score (55.3), which uses a different harness/judge (tool use enabled, GPT-5.6-luna judge, 300K context) and is not directly comparable.
- GPQA Diamond: Independently measured by Artificial Analysis. Exact value from the page's underlying data (0.912121...); AA's on-screen rounded display shows 91%. Not yet present on vals.ai (checked live; vals.ai's newest ZAI model is 'GLM 5.3' released 2026-08-18, a distinct larger non-Flash model, not GLM-5.3-Flash).
- Terminal-Bench 2.1: Independently measured by Artificial Analysis. Exact value from the page's underlying data (0.842697...); AA's on-screen rounded display shows 84%. Consistent with Z.ai's self-reported 84.3 (vendor harness: Claude Code 2.1.207).
- DeepSWE: Reported at the model's "max" effort tier. DeepSWE v1.1 official leaderboard, updated 2026-08-26: 63%±4% Pass@1, avg cost $0.24, 73k output tokens, 123 steps. Closely matches Z.ai's self-reported 63.4 on the same benchmark.
- Toolathlon-Verified: Vendor-reported pass@1 averaged over 3 independent runs via the official Toolathlon evaluation service, per Z.ai's footnote. Not yet on toolathlon.xyz's own leaderboard as of today's live check (its newest ZAI entry is 'GLM 5.2 (max)' at 59.9).
- Agents' Last Exam: Vendor-reported, evaluated using the official ALE protocol with Claude Code harness (effort=max, 1M context, 64K max output, tool search disabled), scored by official ALE evaluators per Z.ai's footnote. Not yet on snorkel.ai's own leaderboard as of today's live check (its newest ZAI entries are GLM-5.1/GLM-5.2).
Notes on the record
GLM-5.3-Flash launched on 2026-08-26 as, in Z.ai's own words, "the first natively multimodal model in the GLM-5 series" — 320B total / 18B active parameters, a hybrid sparse+linear attention architecture with Manifold-Constrained Hyper-Connections (mHC), trained on a 30T-token multimodal corpus (Z.ai blog, as of 2026-08-26). Its spec sheet is contested from day one: the live Hugging Face repository (zai-org/GLM-5.3-Flash) lists an MIT license and real downloadable BF16/F8_E4M3/F32 safetensors alongside a 1,048,576-token (1M) context window, corroborated by OpenRouter and docs.z.ai (as of 2026-08-26), while Artificial Analysis's own same-day model page instead states "Proprietary model," weights not publicly available, text-only input/output, a 400k context window, and an undisclosed parameter count. Because the downloadable weights on Hugging Face are directly verifiable and Artificial Analysis's page is internally inconsistent with them, this site records the model as open-weights, multimodal, and 1M-context — but the disagreement itself is worth knowing before citing either spec on its own.
Coverage is thin at launch: only 6 of the 11 benchmarks this site tracks carry a GLM-5.3-Flash score as of 2026-08-26. Four of those six were measured independently — HLE at 39.9 (standard, no tools), GPQA Diamond at 91.2, and Terminal-Bench 2.1 at 84.3, all via Artificial Analysis, plus DeepSWE v1.1 at 63% (±4% Pass@1) via that benchmark's own official leaderboard, which lists the entry as "glm-5.3-flash [max]." The other two — Toolathlon Verified (78.4) and Agents' Last Exam (26.3) — currently rest on Z.ai's self-reported numbers alone, since neither toolathlon.xyz nor snorkel.ai had posted a GLM-5.3-Flash row as of the live check. Where the two methods overlap they land close together (Terminal-Bench 2.1: 84.27 independent, AA displays 84.3, vs. 84.3 self-reported; DeepSWE: 63 independent vs. 63.4 self-reported), a point in the vendor's favor. But Z.ai's headline "HLE w/ Tools" figure of 55.3 is not one of the six scores above and should not be read against the 39.9 — it comes from a different, more permissive harness (tool use enabled, GPT-5.6-luna judge, 300K context) than the standard no-tools HLE that Artificial Analysis ran.
Z.ai also says it stealth-tested the model before launch under the name "ox-alpha" on OpenCode and OpenRouter, where it "quickly became the most popular model of the week" — a claim OpenRouter's FAQ corroborates on the identity reveal, though that only confirms the product identity, not that the stealth build and the GA release share one checkpoint. LiveBench.ai, re-checked live on 2026-08-26, still carries that build under its stealth name ("ox-alpha-max," 69.2 overall) with no separate GLM-5.3-Flash row; following this site's earlier DeepSeek-V4-Pro checkpoint-mismatch precedent, that score is left off this page rather than credited to the GA model.
Pricing is promotional, not list: $0.075 per 1M input tokens and $0.25 per 1M output tokens (cache read $0.015), in effect through 2026-09-09 24:00 UTC+8 per docs.z.ai — a stated 50% launch discount off the $0.15 in / $0.50 out (cache read $0.03) rate that takes over afterward. OpenRouter confirms the same two-tier structure; Artificial Analysis's displayed price reflects the future list rate rather than today's promo rate, so the two sources aren't in conflict, just quoting different points in time.
Compare with
FAQ
Is GLM-5.3-Flash actually multimodal with a 1M-token context, or is that just Z.ai's marketing?
The live Hugging Face repository backs it up with a 1,048,576-token context window and downloadable, MIT-licensed safetensors weights, corroborated by OpenRouter and docs.z.ai. The multimodal claim comes from Z.ai's own blog post, not from the HF repo's spec fields — Artificial Analysis's own page disagrees on that point too, describing the model as text-only with a 400k-token limit. The two sources simply have not been reconciled yet.
How good is it at coding and agent tasks?
On Terminal-Bench 2.1 it scores 84.3, measured independently by Artificial Analysis (84.27, which AA displays rounded as 84.3) and matched almost exactly by Z.ai's own reported 84.3. On DeepSWE v1.1's official leaderboard the "glm-5.3-flash [max]" entry scores 63% (±4% Pass@1), close to Z.ai's self-reported 63.4. Two more agentic scores, Toolathlon Verified (78.4) and Agents' Last Exam (26.3), rest on Z.ai's own testing alone, with no outside leaderboard to confirm them yet.
Can this model actually be downloaded and self-hosted?
The Hugging Face listing says yes: an MIT license with real BF16, F8_E4M3, and F32 safetensors files ready to pull down, which is the strongest evidence available for calling it open-weights — strong enough that this site records it as open-weights despite Artificial Analysis's same-day page calling it proprietary with no public weights.
How much does it cost to use right now, and will that change?
It launched at a promotional $0.075 per 1M input tokens and $0.25 per 1M output tokens, which Z.ai describes as a 50% launch discount. That rate holds only until 2026-09-09 at 24:00 UTC+8, after which the price doubles to $0.15 in / $0.50 out per Z.ai's own pricing page.
Was this model secretly available before its official announcement?
Z.ai says it ran an anonymous preview called 'ox-alpha' on OpenCode and OpenRouter that "quickly became the most popular model of the week," and OpenRouter's FAQ backs up that this preview and GLM-5.3-Flash are the same product. However, LiveBench.ai still scores that preview under its old name with no benchmark yet confirmed against the officially released version, so its 69.2 overall score is not attributed to GLM-5.3-Flash on this page.