Zhipu AI (Z.ai)Superseded by GLM-5.3

GLM-5.2

GLM-5.2 carries a "superseded" label, but its benchmark record is still the more independently verified of the two: eleven of its twelve tracked scores are independent runs, versus seven of nine for GLM-5.3, the model that replaced it. The one self-reported holdout is Humanity's Last Exam with tool use (54.7); on Terminal-Bench 2.1, Z.ai's 81.0 has given way to vals.ai's independent 67.79.

GLM-5.2 benchmarks and pricing, every number sourced: 12 tracked GLM-5.2 benchmark scores (11 independently run, 1 still resting on a vendor’s own claim), priced at $1.40 per million input tokens and $4.40 per million output.

GLM-5.2 architecture: Mixture-of-Experts with sparse attention; 1M-token context window.

GLM-5.2’s 12 benchmark scores on this page were each verified against their sources between 2026-06-30 and 2026-10-01.

Released
2026-06-16
License
open-weights
Context window
1M tokens
Knowledge cutoff
Not disclosed
Parameters
Not disclosed
Architecture
Mixture-of-Experts with sparse attention

GLM-5.2’s verified record

GLM-5.2’s featured comparison is DeepSeek V4 Pro (0813): 1 trail, 1 tie, and 8 not callable. GLM-5.2 is priced at $1.40/$4.40 per 1M tokens (in/out) vs DeepSeek V4 Pro (0813)’s $1.32/$3.96. Full verdict →

Against the 224 head-to-head comparisons GLM-5.2 shares with other tracked models: 76 real gaps, 15 inside the noise band, and 133 we will not call.

A gap counts for GLM-5.2 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — GLM-5.2 trails on 70 of them.

HLE · no tools Reasoning

±2 is noise
Behind
Qwen3.8-Max −2.0 · Grok 4.7 −2.0 · Step 5 Preview −5.4 · Kimi K3 −5.8 · Gemini 3.1 Pro Preview −5.9 · Gemini 3.8 Flash −6.7 · GPT-6 Sol −6.8 · Muse Spark 1.3 −7.6 · MiMo-V2.6-Pro −8.3 · GPT-5.6 Sol −8.4 · GPT-6.1 Sol −11.8 · GPT-6 Astra −13.6 · Claude Opus 5 −13.8 · Claude Sonnet 5.5 −13.9 · Gemini 4 Argon −16.0 · Claude Fable 5.1 −18.0 · Claude Opus 5.5 −20.3 · 4 superseded: Muse Spark 1.2 −4.4 · Gemini 3.7 Flash −6.8 · Claude Opus 4.8 −7.6 · Claude Fable 5 −14.4
Ahead
GPT-6 Luna +2.6 · Qwen3.8-Flash-Next +3.1 · 1 superseded: DeepSeek V4 Flash (0731) +2.5
Tie
6 models within ±2
Unverified
2 models — vendor-reported on one side

LiveBench Composite score across 7 domains

±2.7 is noise
Behind
Claude Sonnet 5 −2.8 · GLM-5.3 −2.9 · Qwen3.8-Flash-Next −3.0 · Gemini 3.1 Pro Preview −3.8 · DeepSeek V4 Pro (0813) −4.2 · Grok 4.7 −4.2 · Claude Sonnet 5.5 −4.6 · Grok 4.6 −4.8 · Qwen3.8-Max −5.3 · Kimi K3 −6.0 · GPT-6 Sol −6.1 · Claude Opus 5 −6.9 · GPT-5.6 Sol −7.8 · Muse Spark 1.3 −8.4 · GPT-6.1 Sol −8.4 · Claude Opus 5.5 −10.0 · Claude Fable 5.1 −10.2 · 4 superseded: Claude Opus 4.8 −3.0 · Muse Spark 1.2 −4.8 · Gemini 3.7 Flash −5.6 · Claude Fable 5 −9.8
Ahead
Tie
4 models within ±2.7

DeepSWE Long-horizon coding

±9.5 is noise
Behind
Claude Sonnet 5 −10.0 · Qwen3.8-Max −13.0 · GLM-5.3-Flash −19.0 · Grok 4.6 −23.0 · GPT-5.6 Luna −23.0 · Kimi K3 −25.0 · GLM-5.3 −25.0 · GPT-5.6 Sol −29.0 · Claude Opus 5 −30.0 · Gemini 3.8 Flash −30.0 · GPT-6 Astra −30.0 · 4 superseded: Muse Spark 1.2 −11.0 · Claude Opus 4.8 −15.0 · Gemini 3.7 Flash −21.0 · Claude Fable 5 −26.0
Tie
1 model within ±9.5
Unverified
10 models — vendor-reported on one side
Setup-dependent
2 models — scored on a different harness or effort tier

Toolathlon-Verified Multi-tool chores

±9.7 is noise
Behind
Claude Sonnet 5 −11.7 · Kimi K3 −16.6 · 3 superseded: DeepSeek V4 Flash (0731) −10.8 · Muse Spark 1.2 −16.0 · Claude Opus 4.8 −16.3
Tie
1 model within ±9.7
Unverified
10 models — vendor-reported on one side

Agents' Last Exam Professional work

±3.2 is noise
Behind
GPT-6 Luna −4.6 · Kimi K3 −7.9 · GPT-6 Astra −13.8 · Claude Opus 5.5 −17.8 · 1 superseded: Claude Opus 4.8 −6.6
Unverified
6 models — vendor-reported on one side
Setup-dependent
8 models — scored on a different harness or effort tier

Terminal-Bench 2.1 Terminal ops

±10.6 is noise
Behind
Kimi K3 −13.1 · GPT-6 Sol −15.4
Ahead
Tie
3 models within ±10.6
Unverified
1 model — vendor-reported on one side
Setup-dependent
23 models — scored on a different harness or effort tier

ARC-AGI-2 · untiered Compositional visual reasoning

±9.2 is noise
Behind

No verdict for GLM-5.2 anywhere on GPQA Diamond, LiveCodeBench, SWE-bench Verified (saturated); HLE · with tools (nothing independently confirmed on both sides); HMMT (contaminated).

GLM-5.2 API pricing

$1.40 in / $4.40 out per 1M tokens — official pricing

What GLM-5.2 costs per job

GLM-5.2 cost for three reference workloads, computed from its list rates
WorkloadTokens in / outCost
One long chat turn100K / 10K$0.184
A codebase review1,000K / 100K$1.84
A day of agent work10,000K / 1,000K$18.40

Computed from GLM-5.2’s list rates above — cache discounts and batch tiers are not applied.

GLM-5.2 is one of 3 Zhipu AI (Z.ai) models tracked on this site, at these official list prices.

Zhipu AI (Z.ai) model pricing, official list rates
ModelIn / 1MOut / 1M
GLM-5.3-Flash$0.15$0.50
GLM-5.3$1.40$4.40
GLM-5.2(superseded)$1.40$4.40

GLM-5.2 benchmark scores

GLM-5.2 benchmark scores, provenance, and source links
BenchmarkScore
HLE(no tools)[1]
Reasoning · ±2 is noise
HLE(with tools)[2]
Reasoning · ±2 is noise
GPQA Diamondsaturated[3]
Expert science Q&A — not ranked at any gap size
Terminal-Bench 2.1[4]
Terminal ops · ±10.6 is noise
DeepSWE[5]
Long-horizon coding · ±9.5 is noise
Toolathlon-Verified
Multi-tool chores · ±9.7 is noise
Agents' Last Exam[6]
Professional work · ±3.2 is noise
SWE-bench Verifiedsaturated[7]
Bug fixing — not ranked at any gap size
LiveCodeBenchsaturated[8]
Contest coding — not ranked at any gap size
ARC-AGI-2(untiered)[9]
Compositional visual reasoning · ±9.2 is noise
LiveBench[10]
Composite score across 7 domains · ±2.7 is noise
HMMTcontaminated[11]
Competition mathematics — not ranked at any gap size

Who ran these numbers: 11 of 12 independent — vals.ai (4), artificialanalysis.ai (1), deepswe.datacurve.ai (1), toolathlon.xyz (1), snorkel.ai (1), arcprize.org (1), livebench.ai (1), matharena.ai (1); vendor self-reported (1).

  1. HLE: AA's own run, 'GLM-5.2 (max)' (text-only, no tools). Replaces Z.ai's self-reported 40.5.
  2. HLE: Corroborated by https://docs.z.ai/guides/llm/glm-5.2.
  3. GPQA Diamond: vals.ai run, rank 42/133 (updated 2026-08-15). Replaces Z.ai's self-reported 91.2 — a 5.6-point gap between vendor claim and independent run.
  4. Terminal-Bench 2.1: vals.ai's own Terminus 2 run, board row 'GLM 5.2' (67.79 ± 0.99, reasoning effort max, pass@1), from its archived Terminal-Bench 2.1 table. Replaces Z.ai's self-reported 81.0 (Terminus-2 harness, Hugging Face README), 13.2 points higher.
  5. DeepSWE: Independent score (44% ± 2%, $3.92/task).
  6. Agents' Last Exam: Overall pass rate 20.4 (Claude Code, Max; score 41.1; $1,086). Corrects an earlier version of this page, which showed 41.1 — that was the partial-credit score, not the pass rate.
  7. SWE-bench Verified: vals.ai run, rank 14/83, bash-only harness, $0.71/test (updated 2026-08-14).
  8. LiveCodeBench: vals.ai run, rank 88/138 (updated 2026-08-15).
  9. ARC-AGI-2: GLM-5.2's official ARC-AGI-2 leaderboard row, dated 2026-06-13 on arcprize.org. No Low/Medium/High/Max split exists for this row — a single flat score.
  10. LiveBench: Board row "GLM-5.2" on the 2026-06-25 LiveBench release.
  11. HMMT: MathArena's leaderboard row is "GLM 5.2", 92.42% across all 33 HMMT Feb 2026 problems (live-verified 2026-08-25 at matharena.ai/?comp=hmmt--hmmt_feb_2026). Carries MathArena's own contamination-warning flag — "Model was released after competition release" — a disclosed risk that applies to every one of this site's 4 tracked HMMT Feb 2026 rows, GLM 5.2 included, since all 4 tracked models were released after HMMT Feb 2026 took place.

Notes on the record

MIT-licensed weights on Hugging Face and ModelScope. Cached-input price is $0.26/1M, far below the $1.40 shown.

Marked superseded on 2026-08-19 when GLM-5.3's API went live at the identical price — the condition this entry pre-committed to ('keep current until GLM-5.3 is actually available via API'). Still the only open-weights model in the main GLM-5 line until GLM-5.3's promised weights release; the Flash-tier GLM-5.3-Flash shipped MIT-licensed weights on 2026-08-26. It is also the GLM model with the most independent benchmark rows on this site: 11 of its 12 tracked scores are independent runs, versus 7 of 9 for GLM-5.3 (recounted 2026-10-01, when vals.ai's Terminal-Bench 2.1 run replaced Z.ai's self-report; earlier, on 2026-08-27, both models picked up independent rows through late August — GLM-5.3's share grew sharply on 2026-08-20 as Artificial Analysis and vals.ai ran three more of its benchmarks, and its DeepSWE row flipped from vendor self-report to that benchmark's official board on 2026-08-26 — but GLM-5.2 remains ahead on the raw independent-row count). GLM-5.3's weights still have not appeared on Hugging Face; Z.ai said at the August 14 launch to expect them roughly two weeks later (around 2026-08-28), pending a safety review tied to the model's emergent exploit-finding capability.

No parameter count from Z.ai. The GLM-5.2 model card publishes an architecture but no total and no activated figure, so this page leaves both blank rather than borrowing one. HuggingFace's file widget computes 753B from the published weights, and Artificial Analysis lists 40B active; neither is a Z.ai statement, and the two answer different questions than a vendor spec would. Checked 2026-08-28.

Compare with

FAQ

Has GLM-5.2 been independently benchmarked?

Mostly, yes. Eleven of the twelve benchmark scores tracked for GLM-5.2 on this site come from independent runs — Artificial Analysis (HLE, no tools), vals.ai (GPQA Diamond, SWE-bench Verified, LiveCodeBench, Terminal-Bench 2.1), deepswe.datacurve.ai (DeepSWE), Toolathlon, Snorkel AI (Agents' Last Exam), ARC Prize (ARC-AGI-2), LiveBench's own board, and MathArena (HMMT Feb 2026). The one exception is Z.ai's own Humanity's Last Exam with tool use (54.7). Z.ai's Terminal-Bench 2.1 figure, 81.0, was replaced on 2026-10-01 by vals.ai's run, 67.79.

Is GLM-5.2 open source?

Yes. Z.ai published GLM-5.2's weights on Hugging Face and ModelScope under an MIT license. That still matters after the 'superseded' label: GLM-5.3, its replacement, launched proprietary on 2026-08-14, and as of 2026-08-20 its weights still haven't shown up on Hugging Face — Z.ai has said to expect them around 2026-08-28, held for a safety review of the model's exploit-finding capability. Until then, GLM-5.2 is the only model in this pair you can actually download.

Why is GLM-5.2 marked superseded if it's still usable?

Because this entry's own pre-set trigger condition fired, not because the model became unusable. GLM-5.2 was set to stay 'current' only until GLM-5.3 shipped with working API access — that happened on 2026-08-19, at an unchanged $1.40/$4.40 price, so the status flipped that day.

How much does the GLM-5.2 API cost?

$1.40 per million input tokens and $4.40 per million output tokens, per Z.ai's pricing page. Cached input is much cheaper at $0.26 per million tokens — roughly a fifth of the fresh-input rate — which matters for coding workflows that keep resending the same long context.

Is GLM-5.2 good for coding?

On the independently run benchmarks, its record holds up: 82.8 on SWE-bench Verified and 69.5 on LiveCodeBench, both vals.ai, plus 44% on DeepSWE's long-horizon coding tasks. Terminal-Bench 2.1 is the weak spot: Z.ai reported 81.0, but vals.ai's independent Terminus 2 run measured 67.79, 13.2 points lower.

What changed between GLM-5.2 and GLM-5.3?

Z.ai says GLM-5.3 uses the same base model as GLM-5.2, with the gains coming from scaled post-training rather than a new architecture. The API price didn't move — both list at $1.40/$4.40 per million tokens. Z.ai's own launch chart puts GLM-5.3's Terminal-Bench 2.1 score at 88.2 versus GLM-5.2's 81.0, but that comparison is vendor-to-vendor on both ends. Independent runs now exist for both, from different harnesses: vals.ai measured GLM-5.2 at 67.79, and Artificial Analysis measured GLM-5.3 at 83.9, 4.3 points under Z.ai's figure. On vals.ai's own board, which ran both, GLM-5.3 reads 71.54, 3.75 points above GLM-5.2 and inside the 10.6 band.

Further reading

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.