Zhipu AI (Z.ai)
GLM-5.3-Flash
Its downloadable safetensors on Hugging Face (MIT license, 1,048,576-token context) settle the license question the site can verify directly, but the score sheet is less settled: Artificial Analysis's own same-day page instead calls it proprietary, text-only, and 400k-context, and of the nine benchmark scores gathered so far, eight are independently measured — the only Z.ai's own number left is Agents' Last Exam.
GLM-5.3-Flash benchmarks and pricing, every number sourced: 9 tracked GLM-5.3-Flash benchmark scores (8 independently run, 1 still resting on a vendor’s own claim), priced at $0.15 per million input tokens and $0.50 per million output.
GLM-5.3-Flash architecture: Natively multimodal Mixture-of-Experts; 320B total parameters (18B activated per token); 1M-token context window.
- Released
- 2026-08-26
- License
- open-weights
- Context window
- 1M tokens
- Knowledge cutoff
- Not disclosed
- Verified
- sources checked 2026-08-26–2026-10-08
- Parameters
- 320B (18B active)
- Architecture
- Natively multimodal Mixture-of-Experts
GLM-5.3-Flash’s verified record
GLM-5.3-Flash’s most-compared rival is Kimi K3: 2 trails, 3 ties, and 3 not callable across their 8 shared comparisons. GLM-5.3-Flash is priced at $0.15/$0.50 per 1M tokens (in/out) vs Kimi K3’s $3.00/$15.00.
Against the 202 head-to-head comparisons GLM-5.3-Flash shares with other tracked models: 61 real gaps, 47 inside the noise band, and 94 we will not call.
A gap counts for GLM-5.3-Flash only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — GLM-5.3-Flash trails on 56 of them.
HLE · no tools Reasoning
±2 is noiseLiveBench Composite score across 7 domains
±2.7 is noiseARC-AGI-2 · max Compositional visual reasoning
±9.2 is noiseDeepSWE Long-horizon coding
±9.5 is noiseToolathlon-Verified Multi-tool chores
±9.7 is noiseNo verdict for GLM-5.3-Flash anywhere on Terminal-Bench 2.1 ( every independently confirmed comparison inside the noise band); Agents' Last Exam, Terminal-Bench 4.0 ( nothing independently confirmed on both sides); GPQA Diamond ( saturated).
GLM-5.3-Flash API pricing
$0.15 in / $0.50 out per 1M tokens — official pricing source
What GLM-5.3-Flash costs per job
| Workload | Tokens in / out | Cost |
|---|---|---|
| One long chat turn | 100K / 10K | $0.020 |
| A codebase review | 1,000K / 100K | $0.200 |
| A day of agent work | 10,000K / 1,000K | $2.00 |
Computed from GLM-5.3-Flash’s list rates above — cache discounts and batch tiers are not applied.
GLM-5.3-Flash is one of 3 Zhipu AI (Z.ai) models tracked on this site, at these official list prices.
| Model | In / 1M | Out / 1M |
|---|---|---|
| GLM-5.3-Flash | $0.15 | $0.50 |
| GLM-5.3 | $1.40 | $4.40 |
| GLM-5.2(previous version) | $1.40 | $4.40 |
GLM-5.3-Flash benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise | |
| GPQA Diamondsaturated[2] Expert science Q&A — not ranked at any gap size | |
| Terminal-Bench 2.1[3] Terminal ops · ±10.6 is noise | |
| DeepSWE[4] Long-horizon coding · ±9.5 is noise | |
| Toolathlon-Verified[5] Multi-tool chores · ±9.7 is noise | |
| Agents' Last Exam[6] Professional work · ±3.2 is noise | |
| LiveBench[7] Composite score across 7 domains · ±2.7 is noise | |
| ARC-AGI-2(max)[8] Compositional visual reasoning · ±9.2 is noise | |
| Terminal-Bench 4.0[9] Terminal ops · ±12.4 is noise |
Who ran these numbers: 8 of 9 independent — artificialanalysis.ai (4), deepswe.datacurve.ai (1), toolathlon.xyz (1), livebench.ai (1), arcprize.org (1); vendor self-reported (1).
- HLE: Standard HLE (no tools), independently measured by Artificial Analysis. Exact value from the page's underlying data; AA's on-screen rounded display shows 40%. Distinct from Z.ai's self-reported 'HLE w/ Tools' score (55.3), which uses a different harness/judge (tool use enabled, GPT-5.6-luna judge, 300K context) and is not directly comparable.
- GPQA Diamond: Independently measured by Artificial Analysis. Exact value from the page's underlying data; AA's on-screen rounded display shows 91%. Not yet present on vals.ai (checked live; vals.ai's newest ZAI model is 'GLM 5.3' released 2026-08-18, a distinct larger non-Flash model, not GLM-5.3-Flash).
- Terminal-Bench 2.1: Independently measured by Artificial Analysis. Exact value from the page's underlying data; AA's on-screen rounded display shows 84%. Consistent with Z.ai's self-reported 84.3 (vendor harness: Claude Code 2.1.207). Cross-check (2026-10-01): vals.ai's archived Terminus 2 table lists 'GLM 5.3 Flash' at 62.92.
- DeepSWE: Reported at the model's "max" effort tier. DeepSWE v1.1 official leaderboard, updated 2026-08-26: 63%±4% Pass@1, avg cost $0.24, 73k output tokens, 123 steps. Closely matches Z.ai's self-reported 63.4 on the same benchmark.
- Toolathlon-Verified: Value unchanged (78.4) but the citation is upgraded: the figure first appeared as a vendor claim in Z.ai's launch post (averaged over 3 runs via Toolathlon's official evaluation service); toolathlon.xyz's own board now carries row 'GLM 5.3 Flash (max)', pass@1 78.4±1.9 with the site's 'Evaluated by us' badge, entry dated 2026-08-30, so the number is independently run rather than vendor-reported.
- Agents' Last Exam: Vendor-reported, evaluated using the official ALE protocol with Claude Code harness (effort=max, 1M context, 64K max output, tool search disabled), scored by official ALE evaluators per Z.ai's footnote. Not yet on snorkel.ai's own leaderboard as of today's live check (its newest ZAI entries are GLM-5.1/GLM-5.2).
- LiveBench: LiveBench overall for the GA release, read from the live board 2026-08-28 — a NEW row ("GLM-5.3 Flash", $0.031/run) that LiveBench added alongside, not instead of, the stealth-period "ox-alpha-max" row (69.2), which remains unrenamed. The two checkpoints score 2.4 points apart on the same board, which is the cleanest confirmation yet of this site's rule against attributing stealth-period scores to a GA release: had we credited 69.2 to GLM-5.3-Flash on reveal day, we would have published the wrong number.
- ARC-AGI-2: ARC Prize's own leaderboard row 'GLM-5.3-Flash (Max)', dated 2026-08-26: 65.8% at $0.093/task, the top of the model's three populated tiers (High 50.1, Low 27.9).
- Terminal-Bench 4.0: Artificial Analysis board entry 'GLM 5.3 Flash' — mini-swe-agent harness, pass@1 averaged over 3 full passes on all 66 tasks, no effort tier published (raw 0.328283). No vals.ai TB4 row. Official board (grant-funded, Claude Code, none effort, 5-trial pass@1): 35.76 ± 3.54 CI — 2.9 points above this run.
Notes on the record
GLM-5.3-Flash launched on 2026-08-26 as, in Z.ai's own words, "the first natively multimodal model in the GLM-5 series" — 320B total / 18B active parameters, a hybrid sparse+linear attention architecture with Manifold-Constrained Hyper-Connections (mHC), trained on a 30T-token multimodal corpus (Z.ai blog, as of 2026-08-26). Its spec sheet is contested from day one: the live Hugging Face repository (zai-org/GLM-5.3-Flash) lists an MIT license and real downloadable BF16/F8_E4M3/F32 safetensors alongside a 1,048,576-token (1M) context window, corroborated by OpenRouter and docs.z.ai (as of 2026-08-26), while Artificial Analysis's own same-day model page instead states "Proprietary model," weights not publicly available, text-only input/output, a 400k context window, and an undisclosed parameter count. Because the downloadable weights on Hugging Face are directly verifiable and Artificial Analysis's page is internally inconsistent with them, this site records the model as open-weights, multimodal, and 1M-context — but the disagreement itself is worth knowing before citing either spec on its own.
Coverage spans 8 of 14 benchmarks this site tracks (2026-10-08). Six are independently measured — HLE at 39.9 (standard, no tools), GPQA Diamond at 91.2, and Terminal-Bench 2.1 at 84.3, all via Artificial Analysis, plus DeepSWE v1.1 at 63% (±4% Pass@1) via its own official leaderboard, which lists the entry as "glm-5.3-flash [max]," and Terminal-Bench 4.0 at 32.83 via vals.ai (2026-10-08) The other two — Toolathlon Verified (78.4) and Agents' Last Exam (26.3) — rest on Z.ai's self-reported numberss alone, since neither toolathlon.xyz nor snorkel.ai had posted a GLM-5.3-Flash row as of the live check. Where the two methods overlap they land close together (Terminal-Bench 2.1: 84.27 independent, AA displays 84.3, vs. 84.3 self-reported; DeepSWE: 63 independent vs. 63.4 self-reported), a point in the vendor's favor. But Z.ai's headline "HLE w/ Tools" figure of 55.3 is not one of the six scores above and should not be read against the 39.9 — it comes from a different, more permissive harness (tool use enabled, GPT-5.6-luna judge, 300K context) than the standard no-tools HLE that Artificial Analysis ran.
Z.ai also says it stealth-tested the model before launch under the name "ox-alpha" on OpenCode and OpenRouter, where it "quickly became the most popular model of the week" — a claim OpenRouter's FAQ corroborates on the identity reveal, though that only confirms the product identity, not that the stealth build and the GA release share one checkpoint. LiveBench settled the checkpoint question itself on 2026-08-28: it added a separate "GLM-5.3 Flash" row scoring 71.6 overall while keeping the stealth-period "ox-alpha-max" row (69.2) unrenamed beside it. The two builds score 2.4 points apart on the same board — direct confirmation that this site's earlier refusal to credit the stealth score to the GA release (following the DeepSeek-V4-Pro checkpoint-mismatch precedent) was the right call: crediting 69.2 on reveal day would have published the wrong number. The GA row's 71.6 is now tracked below; the ox-alpha-max 69.2 remains excluded.
Pricing is now list: $0.15 per 1M input tokens and $0.50 per 1M output tokens (cache read $0.03) per docs.z.ai. It launched 2026-08-26 at a promotional $0.075 / $0.25 (cache read $0.015) — a stated 50% launch discount that ran through 2026-09-09 24:00 UTC+8 and expired as scheduled; OpenRouter's price feed, rechecked 2026-09-22, now returns the list rate. Artificial Analysis displayed the list rate even during the promo, so the sources weren't in conflict — just quoting different points in time.
Compare with
FAQ
Is GLM-5.3-Flash actually multimodal with a 1M-token context, or is that just Z.ai's marketing?
The live Hugging Face repository backs it up with a 1,048,576-token context window and downloadable, MIT-licensed safetensors weights, corroborated by OpenRouter and docs.z.ai. The multimodal claim comes from Z.ai's own blog post, not from the HF repo's spec fields — Artificial Analysis's own page disagrees on that point too, describing the model as text-only with a 400k-token limit. The two sources simply have not been reconciled yet.
How good is it at coding and agent tasks?
On Terminal-Bench 2.1 it scores 84.3, measured independently by Artificial Analysis (84.27, which AA displays rounded as 84.3) and matched almost exactly by Z.ai's own reported 84.3. On DeepSWE v1.1's official leaderboard the "glm-5.3-flash [max]" entry scores 63% (±4% Pass@1), close to Z.ai's self-reported 63.4. Agents' Last Exam (26.3) still rests on Z.ai's own testing alone; Toolathlon Verified (78.4) no longer does — the official board's own run recorded it as an independent result on 2026-10-03.
Can this model actually be downloaded and self-hosted?
The Hugging Face listing says yes: an MIT license with real BF16, F8_E4M3, and F32 safetensors files ready to pull down, which is the strongest evidence available for calling it open-weights — strong enough that this site records it as open-weights despite Artificial Analysis's same-day page calling it proprietary with no public weights.
How much does it cost to use right now, and will that change?
The list rate is $0.15 per 1M input tokens and $0.50 per 1M output tokens per Z.ai's own pricing page. It launched at a promotional $0.075 / $0.25 — a 50% launch discount — but that rate expired on 2026-09-09 at 24:00 UTC+8, and OpenRouter's price feed confirmed the move to the list rate when rechecked on 2026-09-22.
Was this model secretly available before its official announcement?
Z.ai says it ran an anonymous preview called 'ox-alpha' on OpenCode and OpenRouter that "quickly became the most popular model of the week," and OpenRouter's FAQ backs up that this preview and GLM-5.3-Flash are the same product. The scoreboard now shows why product identity and checkpoint identity are different things: on 2026-08-28 LiveBench added a separate GLM-5.3 Flash row at 71.6 overall while the stealth-period ox-alpha-max row still sits at 69.2 beside it — the GA build and the preview build score 2.4 points apart on the same board. This page tracks the 71.6 and has never credited the 69.2.
What is GLM-5.3-Flash's best score on a current benchmark?
Terminal-Bench 2.1 at 84.3 (2026-08-26), followed by Toolathlon-Verified at 78.4 and ARC-AGI-2 at 65.8 on a max-effort run. Its single biggest printed figure, GPQA Diamond 91.2, lives on a saturated board and gets no ranking weight here.
Further reading
- GPQA Diamond leaderboard 2026 — GLM-5.3-Flash is one of the 23 models it compares.
- GLM-5.3-Flash vs DeepSeek Flash vs Qwen3.8 — GLM-5.3-Flash is one of the 3 models it compares.
- Models with 10M token context windows 2026 — GLM-5.3-Flash is one of the 37 models it compares.