Zhipu AI (Z.ai)
GLM-5.3
The launch chart claims frontier coding at an unchanged mid-tier price. The price is real — but seven of the nine benchmark results on this page have been independently run so far (Terminal-Bench 2.1, Humanity's Last Exam without tools, GPQA Diamond, LiveCodeBench, SWE-bench Verified, LiveBench, and now DeepSWE), though GPQA Diamond and SWE-bench Verified are both graded saturated here, so those two settle nothing; only the Toolathlon-Verified coding score is still Z.ai's own run.
GLM-5.3 benchmarks and pricing, every number sourced: 9 tracked GLM-5.3 benchmark scores (7 independently run, 2 still resting on a vendor’s own claim), priced at $1.40 per million input tokens and $4.40 per million output.
GLM-5.3’s 9 benchmark scores on this page were each verified against their sources between 2026-08-19 and 2026-08-26.
Version history: succeeded GLM-5.2 (2026-06-16).
- Released
- 2026-08-14
- License
- proprietary
- Context window
- 1M tokens
- Knowledge cutoff
- Not disclosed
- Parameters
- Not disclosed
- Architecture
- Not disclosed
GLM-5.3’s verified record
GLM-5.3’s most-compared rival is Kimi K3: 2 trails, 1 tie, and 6 not callable across their 9 shared comparisons. GLM-5.3 is priced at $1.40/$4.40 per 1M tokens (in/out) vs Kimi K3’s $3.00/$15.00.
Against the 200 head-to-head comparisons GLM-5.3 shares with other tracked models: 43 real gaps, 41 inside the noise band, and 116 we will not call.
A gap counts for GLM-5.3 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — GLM-5.3 trails on 28 of them.
HLE · no tools Reasoning
±2 is noiseLiveBench Composite score across 7 domains
±2.7 is noiseDeepSWE Long-horizon coding
±9.5 is noiseNo verdict for GLM-5.3 anywhere on Terminal-Bench 2.1 (every independently confirmed comparison inside the noise band); GPQA Diamond, LiveCodeBench, SWE-bench Verified (saturated); HLE · with tools, Toolathlon-Verified (nothing independently confirmed on both sides).
- No real gap yet in any of GLM-5.3’s agentic comparisons — the independently confirmed ones all sit inside the noise band.
What changed from GLM-5.2 to GLM-5.3
The 8 benchmarks both models have been scored on, using the same variant each time. A raw GLM-5.3 gain is not a real gain until it clears that benchmark’s own noise band, so each row below carries the verdict and not just the arithmetic.
| Benchmark | GLM-5.2 | GLM-5.3 | Change | Verdict |
|---|---|---|---|---|
| HLE | 41.1 | 42.3 | +1.2 | Tie |
| Terminal-Bench 2.1 | 67.79 | 83.9 | +16.1 | Setup-dependent |
| DeepSWE | 44 | 69 | +25.0 | Real gap |
| Toolathlon-Verified | 59.9 | 73 | +13.1 | Unverified |
| SWE-bench Verified | 82.8 | 95.4 | +12.6 | Tainted |
| GPQA Diamond | 85.61 | 91.7 | +6.1 | Tainted |
| LiveCodeBench | 69.5 | 80.53 | +11.0 | Tainted |
| LiveBench | 73.2 | 76.1 | +2.9 | Real gap |
GLM-5.3 API pricing
$1.40 in / $4.40 out per 1M tokens — official pricing
What GLM-5.3 costs per job
| Workload | Tokens in / out | Cost |
|---|---|---|
| One long chat turn | 100K / 10K | $0.184 |
| A codebase review | 1,000K / 100K | $1.84 |
| A day of agent work | 10,000K / 1,000K | $18.40 |
Computed from GLM-5.3’s list rates above — cache discounts and batch tiers are not applied.
GLM-5.3 is one of 3 Zhipu AI (Z.ai) models tracked on this site, at these official list prices.
| Model | In / 1M | Out / 1M |
|---|---|---|
| GLM-5.3-Flash | $0.15 | $0.50 |
| GLM-5.3 | $1.40 | $4.40 |
| GLM-5.2(superseded) | $1.40 | $4.40 |
GLM-5.3 benchmark scores
| Benchmark | Score |
|---|---|
| Terminal-Bench 2.1[1] Terminal ops · ±10.6 is noise | |
| DeepSWE[2] Long-horizon coding · ±9.5 is noise | |
| Toolathlon-Verified[3] Multi-tool chores · ±9.7 is noise | |
| HLE(with tools)[4] Reasoning · ±2 is noise | |
| HLE(no tools)[5] Reasoning · ±2 is noise | |
| GPQA Diamondsaturated[6] Expert science Q&A — not ranked at any gap size | |
| LiveCodeBenchsaturated[7] Contest coding — not ranked at any gap size | |
| SWE-bench Verifiedsaturated[8] Bug fixing — not ranked at any gap size | |
| LiveBench[9] Composite score across 7 domains · ±2.7 is noise |
Who ran these numbers: 7 of 9 independent — artificialanalysis.ai (3), vals.ai (2), deepswe.datacurve.ai (1), livebench.ai (1); vendor self-reported (2).
- Terminal-Bench 2.1: AA's own run of 'GLM-5.3 (max)' (Terminus 2 harness, e2b sandbox, pass@1 avg of 3 repeats). Replaces Z.ai's self-reported 88.2 from the launch chart. The 4.3-point gap sits inside this benchmark's 10.6-point noise band, so it is not evidence of a bad number — but across the 20 vendor-vs-independent pairs measured on this site, every gap of 4 points or more has run vendor-high, while the largest gap the other way is 3.4. (That count predates 2026-10-01, when vals.ai rows replaced seven more vendor figures, all of them vendor-high by 7.4 to 30.3 points.) Cross-check (2026-10-01): vals.ai's archived Terminus 2 table lists 'GLM 5.3' at 71.54.
- DeepSWE: deepswe.datacurve.ai v1.1 leaderboard, rank #4 of 18 tracked models as of 2026-08-26 (69% ± 3%, avg cost $3.99/task, 80k output tokens, 124 steps, max effort). This is now on the board — Z.ai's own launch-chart self-report was 66.9%, so the independent run actually landed 2.1 points ABOVE the vendor's own number, the reverse of this site's more common vendor-favorable pattern; the gap sits well inside this benchmark's ±9.5 cross-model noise band either way, so it isn't evidence of anything more than normal run-to-run variance.
- Toolathlon-Verified: Z.ai says results came via Toolathlon's official evaluation service (pass@1, avg of 3 runs), but the number is vendor-published — flips to independent when it appears on toolathlon.xyz's own board. Sits above GLM-5.2's independent 59.9 but below Kimi K3's independent 76.5.
- HLE: Z.ai's own run: temp 1.0, top_p 0.95, 300K max context with context management, GPT-5.6-luna (medium) as judge. Every model's with_tools HLE number on this site is vendor-reported — no independent board runs this variant.
- HLE: AA's own run, 'GLM-5.3 (max)' (text-only, no tools). Z.ai never published a no-tools HLE figure — only the 62.5 with-tools number — so this is the first look at the model without scaffolding.
- GPQA Diamond: AA's own run of 'GLM-5.3 (max)'. Every other GPQA row on this site comes from vals.ai; this one carries a different harness (artificial-analysis) recorded for completeness, though it makes no difference to the verdict here: GPQA Diamond is graded saturated on this site (2026-08-20), so compareScores() marks any pairing on this benchmark ⊘ Tainted before a harness check is ever reached — this row exists for the record, not for ranking.
- LiveCodeBench: vals.ai run, rank 66/140 (board updated 2026-08-19) — required clicking past the top-16 default view to surface. Cost shown alongside ($1.4/$4.4 in/out) matches this site's own tracked GLM-5.3 pricing, cross-confirming the row is the current 0813-era model, not an older GLM revision.
- SWE-bench Verified: vals.ai run, rank 6/86 shown, Mini-SWE-agent harness (board updated 2026-08-19). SWE-bench Verified is graded saturated on this site, so this row completes GLM-5.3's tracked record here without settling any head-to-head comparison.
- LiveBench: Board row "GLM-5.3" on the 2026-06-25 LiveBench release — missed in this site's 2026-08-24 scrape (model released 2026-08-14, before that scrape date) and backfilled 2026-08-25 after a data-QA review flagged the gap. Distinct row from GLM-5.2 (73.2); no version confusion.
Notes on the record
API pricing went live 2026-08-19: $1.40 in / $4.40 out per 1M tokens — identical to GLM-5.2, with cached input at $0.26. Same base model as GLM-5.2; Z.ai says every gain comes from scaled post-training. Reasoning is always on (low / high / max effort — thinking can no longer be disabled, a breaking API change from GLM-5.2).
Artificial Analysis has since run the model independently on Terminal-Bench 2.1 (83.9, 4.3 points below Z.ai's launch chart of 88.2 — inside that benchmark's noise band), Humanity's Last Exam without tools, and GPQA Diamond. vals.ai has independently run it on LiveCodeBench and SWE-bench Verified, though both are graded saturated here. deepswe.datacurve.ai's board picked the model up on or before 2026-08-20 at 69% (rank #4/18) — 2.1 points above Z.ai's own 66.9% launch-chart figure, inside the benchmark's noise band either way. Worth knowing how the Terminal-Bench gap compares: see the FAQ below on GLM-5.3's coding scores for the vendor-vs-independent pattern across tracked pairs on this site. Only the Toolathlon number and the with-tools HLE score are still Z.ai's own. One deliberate omission remains: Z.ai's Agents' Last Exam figure (28.5) is the 105-task ALE-CLI split, not the 1,000-item overall pass rate this site tracks. The GPQA Diamond row is tagged with a different evaluation harness (Artificial Analysis) than the vals.ai-sourced rows most of that column uses — but the row is moot for ranking either way, since GPQA Diamond is graded saturated on this site.
Open weights are still promised within two weeks of launch, held for safety evaluation of the model's emergent exploit-finding capability (84.5% on CyberGym, vendor-run).
Compare with
FAQ
Is GLM-5.3 open source?
Not yet. GLM-5.3 launched proprietary on 2026-08-14 and Artificial Analysis still lists it as closed-weights. Z.ai has promised open weights within two weeks of launch, held back for safety evaluation of the model's emergent exploit-finding capability. Its predecessor GLM-5.2 already ships MIT-licensed open weights, so the family has form here.
How much does the GLM-5.3 API cost?
$1.40 per million input tokens and $4.40 per million output tokens — identical to GLM-5.2 — with cached input at $0.26. API pricing went live on 2026-08-19; before that the model was only reachable through the GLM Coding Plan subscription and ZCode.
Is GLM-5.3 good for coding?
Z.ai's launch chart says frontier-level: 88.2 on Terminal-Bench 2.1 and 66.9 on DeepSWE, with open-source SOTA claims. Artificial Analysis has since run Terminal-Bench 2.1 independently and got 83.9 — 4.3 points lower. That sits inside the benchmark's 10.6-point noise band, so it is not a caught exaggeration. It is worth context though: across the vendor-versus-independent pairs this site has measured, differences under 4 points fall either way about evenly, while every gap of 4 points or more has favored the vendor. DeepSWE has since been independently confirmed too (deepswe.datacurve.ai, 69% vs. Z.ai's own 66.9% — inside the noise band either way); only the Toolathlon number here is still Z.ai's alone.
What is the difference between GLM-5.3 and GLM-5.2?
Same base model — Z.ai says every gain comes from scaled post-training. The API price is unchanged, but thinking is now always-on (low / high / max reasoning effort) and can no longer be disabled, a breaking API change. On Z.ai's own numbers, Terminal-Bench 2.1 moves from 81.0 to 88.2 and DeepSWE from 46.2 to 66.9 — both are Z.ai's own before/after figures from the same launch comparison. Independently, this site's Terminal-Bench 2.1 rows are Artificial Analysis's 83.9 for GLM-5.3 and vals.ai's 67.79 for GLM-5.2, different harnesses; on vals.ai's own board the two read 71.54 and 67.79, a tie. For DeepSWE specifically, this site's own tracked score for GLM-5.2 is 44 (deepswe.datacurve.ai's independent v1.1 leaderboard, max effort), close to but not the same as Z.ai's self-reported 46.2 for the same model and benchmark version.
What is GLM-5.3's context window?
1 million tokens, with a maximum output length of 128K tokens. The knowledge cutoff has not been disclosed.
Why is there no Agents' Last Exam score for GLM-5.3 here?
Because the 28.5 in Z.ai's launch table is the 105-task ALE-CLI split, not the 1,000-item overall pass rate this site tracks. Mixing the two metrics is exactly how benchmark tables mislead — this site refuses numbers that don't match the tracked metric.
Has GLM-5.3 been independently benchmarked?
Yes, on seven of the nine benchmarks tracked for it, as of 2026-08-26. Artificial Analysis has run it independently on Terminal-Bench 2.1 (83.9), a no-tools Humanity's Last Exam score (42.3), and GPQA Diamond (91.7). vals.ai has independently run it on LiveCodeBench (80.53) and SWE-bench Verified (95.40). LiveBench has run it on its own harness too, at 76.1 (Overall). deepswe.datacurve.ai has independently run DeepSWE at 69% (rank #4/18), edging out Z.ai's own 66.9% self-report by a margin well inside the benchmark's noise band. Toolathlon-Verified and with-tools HLE remain Z.ai's own. GPQA Diamond and SWE-bench Verified are both graded saturated on this site, so no pairing on either ever resolves to Real or Tie regardless of the numbers involved — those two rows exist for the record, not for ranking. The GPQA Diamond row is also tagged with a different evaluation harness than the vals.ai-sourced rows elsewhere in that column, recorded for completeness, though saturation alone already rules out a verdict there.
Further reading
- GPQA Diamond leaderboard 2026 — GLM-5.3 is one of the 23 models it compares.
- Models with 10M token context windows 2026 — GLM-5.3 is one of the 34 models it compares.