Zhipu AI (Z.ai)Superseded by GLM-5.3
GLM-5.2
GLM-5.2 carries a "superseded" label, but its benchmark record is still the more independently verified of the two: eleven of its twelve tracked scores are independent runs, versus seven of nine for GLM-5.3, the model that replaced it. The one self-reported holdout is Humanity's Last Exam with tool use (54.7); on Terminal-Bench 2.1, Z.ai's 81.0 has given way to vals.ai's independent 67.79.
GLM-5.2 benchmarks and pricing, every number sourced: 12 tracked GLM-5.2 benchmark scores (11 independently run, 1 still resting on a vendor’s own claim), priced at $1.40 per million input tokens and $4.40 per million output.
GLM-5.2 architecture: Mixture-of-Experts with sparse attention; 1M-token context window.
GLM-5.2’s 12 benchmark scores on this page were each verified against their sources between 2026-06-30 and 2026-10-01.
- Released
- 2026-06-16
- License
- open-weights
- Context window
- 1M tokens
- Knowledge cutoff
- Not disclosed
- Parameters
- Not disclosed
- Architecture
- Mixture-of-Experts with sparse attention
GLM-5.2’s verified record
GLM-5.2’s featured comparison is DeepSeek V4 Pro (0813): 1 trail, 1 tie, and 8 not callable. GLM-5.2 is priced at $1.40/$4.40 per 1M tokens (in/out) vs DeepSeek V4 Pro (0813)’s $1.32/$3.96. Full verdict →
Against the 224 head-to-head comparisons GLM-5.2 shares with other tracked models: 76 real gaps, 15 inside the noise band, and 133 we will not call.
A gap counts for GLM-5.2 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — GLM-5.2 trails on 70 of them.
HLE · no tools Reasoning
±2 is noiseLiveBench Composite score across 7 domains
±2.7 is noiseDeepSWE Long-horizon coding
±9.5 is noiseToolathlon-Verified Multi-tool chores
±9.7 is noiseAgents' Last Exam Professional work
±3.2 is noiseTerminal-Bench 2.1 Terminal ops
±10.6 is noiseARC-AGI-2 · untiered Compositional visual reasoning
±9.2 is noiseNo verdict for GLM-5.2 anywhere on GPQA Diamond, LiveCodeBench, SWE-bench Verified (saturated); HLE · with tools (nothing independently confirmed on both sides); HMMT (contaminated).
GLM-5.2 API pricing
$1.40 in / $4.40 out per 1M tokens — official pricing
What GLM-5.2 costs per job
| Workload | Tokens in / out | Cost |
|---|---|---|
| One long chat turn | 100K / 10K | $0.184 |
| A codebase review | 1,000K / 100K | $1.84 |
| A day of agent work | 10,000K / 1,000K | $18.40 |
Computed from GLM-5.2’s list rates above — cache discounts and batch tiers are not applied.
GLM-5.2 is one of 3 Zhipu AI (Z.ai) models tracked on this site, at these official list prices.
| Model | In / 1M | Out / 1M |
|---|---|---|
| GLM-5.3-Flash | $0.15 | $0.50 |
| GLM-5.3 | $1.40 | $4.40 |
| GLM-5.2(superseded) | $1.40 | $4.40 |
GLM-5.2 benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise | |
| HLE(with tools)[2] Reasoning · ±2 is noise | |
| GPQA Diamondsaturated[3] Expert science Q&A — not ranked at any gap size | |
| Terminal-Bench 2.1[4] Terminal ops · ±10.6 is noise | |
| DeepSWE[5] Long-horizon coding · ±9.5 is noise | |
| Toolathlon-Verified Multi-tool chores · ±9.7 is noise | |
| Agents' Last Exam[6] Professional work · ±3.2 is noise | |
| SWE-bench Verifiedsaturated[7] Bug fixing — not ranked at any gap size | |
| LiveCodeBenchsaturated[8] Contest coding — not ranked at any gap size | |
| ARC-AGI-2(untiered)[9] Compositional visual reasoning · ±9.2 is noise | |
| LiveBench[10] Composite score across 7 domains · ±2.7 is noise | |
| HMMTcontaminated[11] Competition mathematics — not ranked at any gap size |
Who ran these numbers: 11 of 12 independent — vals.ai (4), artificialanalysis.ai (1), deepswe.datacurve.ai (1), toolathlon.xyz (1), snorkel.ai (1), arcprize.org (1), livebench.ai (1), matharena.ai (1); vendor self-reported (1).
- HLE: AA's own run, 'GLM-5.2 (max)' (text-only, no tools). Replaces Z.ai's self-reported 40.5.
- HLE: Corroborated by https://docs.z.ai/guides/llm/glm-5.2.
- GPQA Diamond: vals.ai run, rank 42/133 (updated 2026-08-15). Replaces Z.ai's self-reported 91.2 — a 5.6-point gap between vendor claim and independent run.
- Terminal-Bench 2.1: vals.ai's own Terminus 2 run, board row 'GLM 5.2' (67.79 ± 0.99, reasoning effort max, pass@1), from its archived Terminal-Bench 2.1 table. Replaces Z.ai's self-reported 81.0 (Terminus-2 harness, Hugging Face README), 13.2 points higher.
- DeepSWE: Independent score (44% ± 2%, $3.92/task).
- Agents' Last Exam: Overall pass rate 20.4 (Claude Code, Max; score 41.1; $1,086). Corrects an earlier version of this page, which showed 41.1 — that was the partial-credit score, not the pass rate.
- SWE-bench Verified: vals.ai run, rank 14/83, bash-only harness, $0.71/test (updated 2026-08-14).
- LiveCodeBench: vals.ai run, rank 88/138 (updated 2026-08-15).
- ARC-AGI-2: GLM-5.2's official ARC-AGI-2 leaderboard row, dated 2026-06-13 on arcprize.org. No Low/Medium/High/Max split exists for this row — a single flat score.
- LiveBench: Board row "GLM-5.2" on the 2026-06-25 LiveBench release.
- HMMT: MathArena's leaderboard row is "GLM 5.2", 92.42% across all 33 HMMT Feb 2026 problems (live-verified 2026-08-25 at matharena.ai/?comp=hmmt--hmmt_feb_2026). Carries MathArena's own contamination-warning flag — "Model was released after competition release" — a disclosed risk that applies to every one of this site's 4 tracked HMMT Feb 2026 rows, GLM 5.2 included, since all 4 tracked models were released after HMMT Feb 2026 took place.
Notes on the record
MIT-licensed weights on Hugging Face and ModelScope. Cached-input price is $0.26/1M, far below the $1.40 shown.
Marked superseded on 2026-08-19 when GLM-5.3's API went live at the identical price — the condition this entry pre-committed to ('keep current until GLM-5.3 is actually available via API'). Still the only open-weights model in the main GLM-5 line until GLM-5.3's promised weights release; the Flash-tier GLM-5.3-Flash shipped MIT-licensed weights on 2026-08-26. It is also the GLM model with the most independent benchmark rows on this site: 11 of its 12 tracked scores are independent runs, versus 7 of 9 for GLM-5.3 (recounted 2026-10-01, when vals.ai's Terminal-Bench 2.1 run replaced Z.ai's self-report; earlier, on 2026-08-27, both models picked up independent rows through late August — GLM-5.3's share grew sharply on 2026-08-20 as Artificial Analysis and vals.ai ran three more of its benchmarks, and its DeepSWE row flipped from vendor self-report to that benchmark's official board on 2026-08-26 — but GLM-5.2 remains ahead on the raw independent-row count). GLM-5.3's weights still have not appeared on Hugging Face; Z.ai said at the August 14 launch to expect them roughly two weeks later (around 2026-08-28), pending a safety review tied to the model's emergent exploit-finding capability.
No parameter count from Z.ai. The GLM-5.2 model card publishes an architecture but no total and no activated figure, so this page leaves both blank rather than borrowing one. HuggingFace's file widget computes 753B from the published weights, and Artificial Analysis lists 40B active; neither is a Z.ai statement, and the two answer different questions than a vendor spec would. Checked 2026-08-28.
Compare with
FAQ
Has GLM-5.2 been independently benchmarked?
Mostly, yes. Eleven of the twelve benchmark scores tracked for GLM-5.2 on this site come from independent runs — Artificial Analysis (HLE, no tools), vals.ai (GPQA Diamond, SWE-bench Verified, LiveCodeBench, Terminal-Bench 2.1), deepswe.datacurve.ai (DeepSWE), Toolathlon, Snorkel AI (Agents' Last Exam), ARC Prize (ARC-AGI-2), LiveBench's own board, and MathArena (HMMT Feb 2026). The one exception is Z.ai's own Humanity's Last Exam with tool use (54.7). Z.ai's Terminal-Bench 2.1 figure, 81.0, was replaced on 2026-10-01 by vals.ai's run, 67.79.
Is GLM-5.2 open source?
Yes. Z.ai published GLM-5.2's weights on Hugging Face and ModelScope under an MIT license. That still matters after the 'superseded' label: GLM-5.3, its replacement, launched proprietary on 2026-08-14, and as of 2026-08-20 its weights still haven't shown up on Hugging Face — Z.ai has said to expect them around 2026-08-28, held for a safety review of the model's exploit-finding capability. Until then, GLM-5.2 is the only model in this pair you can actually download.
Why is GLM-5.2 marked superseded if it's still usable?
Because this entry's own pre-set trigger condition fired, not because the model became unusable. GLM-5.2 was set to stay 'current' only until GLM-5.3 shipped with working API access — that happened on 2026-08-19, at an unchanged $1.40/$4.40 price, so the status flipped that day.
How much does the GLM-5.2 API cost?
$1.40 per million input tokens and $4.40 per million output tokens, per Z.ai's pricing page. Cached input is much cheaper at $0.26 per million tokens — roughly a fifth of the fresh-input rate — which matters for coding workflows that keep resending the same long context.
Is GLM-5.2 good for coding?
On the independently run benchmarks, its record holds up: 82.8 on SWE-bench Verified and 69.5 on LiveCodeBench, both vals.ai, plus 44% on DeepSWE's long-horizon coding tasks. Terminal-Bench 2.1 is the weak spot: Z.ai reported 81.0, but vals.ai's independent Terminus 2 run measured 67.79, 13.2 points lower.
What changed between GLM-5.2 and GLM-5.3?
Z.ai says GLM-5.3 uses the same base model as GLM-5.2, with the gains coming from scaled post-training rather than a new architecture. The API price didn't move — both list at $1.40/$4.40 per million tokens. Z.ai's own launch chart puts GLM-5.3's Terminal-Bench 2.1 score at 88.2 versus GLM-5.2's 81.0, but that comparison is vendor-to-vendor on both ends. Independent runs now exist for both, from different harnesses: vals.ai measured GLM-5.2 at 67.79, and Artificial Analysis measured GLM-5.3 at 83.9, 4.3 points under Z.ai's figure. On vals.ai's own board, which ran both, GLM-5.3 reads 71.54, 3.75 points above GLM-5.2 and inside the 10.6 band.
Further reading
- GPQA Diamond leaderboard 2026 — GLM-5.2 is one of the 23 models it compares.
- Models with 10M token context windows 2026 — GLM-5.2 is one of the 35 models it compares.