Three cheap Chinese models, only one fully independent scorecard

August 27, 2026

GLM-5.3-Flash, Qwen3.8-Flash-Next, and DeepSeek V4 Flash are three open-weight (GLM-5.3-Flash's status is disputed — more below) budget-tier models from three different Chinese labs, all released within about a month of each other. Between them, this site tracks 18 benchmark line items. Every one of DeepSeek's eight comes from an independent evaluator. GLM-5.3-Flash's six split four independent and two self-reported. Qwen3.8-Flash-Next's four are all self-reported — and as of 2026-08-26, this site could not find the model on a single one of eight independent leaderboards it checked. The scores below are real. Who ran them is the actual comparison.

The split that doesn't average out

Independent-score ratio: DeepSeek V4 Flash 8 of 8 (100%), GLM-5.3-Flash 4 of 6 (67%), Qwen3.8-Flash-Next 0 of 4 (0%).

DeepSeek is the elder of the three by about four weeks — it shipped 2026-07-31, while GLM-5.3-Flash and Qwen3.8-Flash-Next both landed 2026-08-26. That head start doesn't explain the gap on its own: independent boards have had a full month to score DeepSeek, but they've also had roughly six days to score GLM-5.3-Flash under its former alias (more on that below), and — counting from when its downloadable weights actually appeared, not its official reveal date — roughly two days for Qwen3.8-Flash-Next, whose Hugging Face repo went up 2026-08-24.

Independent doesn't mean better. It means a party with no stake in the outcome ran the eval and published the number. A self-reported score can still be accurate — Qwen's four numbers might hold up perfectly once outside evaluators get to them. But right now, nobody outside Alibaba has checked. That's the entire content of the "0 of 4."

Three benchmarks, three provenance tags

Only three of the eleven benchmarks this site tracks have a published number for all three models:

BenchmarkDeepSeek V4 FlashGLM-5.3-FlashQwen3.8-Flash-Next
HLE (no tools / thinking*)38.6 — independent39.9 — independent35.9* — self-reported
DeepSWE v1.153.0 — independent63% — independent (official board, "max" tier)58.7 — self-reported
Toolathlon-Verified70.7 — independent78.4 — self-reported73.5 — self-reported (Pass@1)

*Three caveats before reading anything into those rows. Qwen's HLE score runs in "thinking mode," the model's default, graded by GPT-4o as judge — a different setup from the no-tools variant DeepSeek and GLM-5.3-Flash both report, so it isn't a clean three-way match even before provenance enters the picture. GLM-5.3-Flash's DeepSWE figure is pulled from the official board's "max" tier; DeepSeek's row doesn't specify a tier, so a 10-point gap between two independently-sourced numbers may not be as clean as it looks. And only Qwen's Toolathlon-Verified score is explicitly labeled Pass@1 — DeepSeek's and GLM-5.3-Flash's aren't, so there's no confirmation all three rows measure the same pass criterion.

Narrow the table to cells this site can actually check against an outside source, and two rows carry an independent number on both sides: HLE and DeepSWE v1.1, for DeepSeek and GLM-5.3-Flash. Only one of the two clears this site's noise band. On HLE, GLM-5.3-Flash's 39.9 against DeepSeek's 38.6 is a 1.3-point gap on a 2,500-question benchmark whose meaningful-gap threshold is 2.0 points — inside the band, so it logs as a tie, not a lead, no matter which name is on top. On DeepSWE v1.1 the 10-point gap does clear that benchmark's 9.5-point threshold, but only just, and the tier caveat above still applies — so treat it as provenance-clean, not necessarily measurement-clean. Toolathlon-Verified has exactly one independent number in its row — DeepSeek's 70.7 — so there's nothing independent to weigh it against yet. None of this is a verdict; it's just what's left once the self-reported numbers are set aside.

GLM-5.3-Flash and DeepSeek V4 Flash also share a fourth benchmark, GPQA Diamond — 91.2 versus 89.9, both independent. This site grades GPQA Diamond saturated, so per its own methodology that gap never drives a verdict, regardless of who ran it. GLM-5.3-Flash and Qwen3.8-Flash-Next share one of their own, Agents' Last Exam — 26.3 versus 24.3, both self-reported. The same pass-criterion gap applies here as above: Qwen's figure is labeled Pass@1, GLM-5.3-Flash's isn't, so it's not confirmed both numbers measure the same thing. With provenance on neither side either way, that pair logs as unverified rather than a real gap or a tie.

Ox Alpha, unmasked

GLM-5.3-Flash's launch carries a second story: it's the model that spent roughly six days running anonymously on OpenRouter and OpenCode as "ox-alpha," starting around 2026-08-20. This site's own search data showed real reader demand for "ox alpha" benchmark queries before Zhipu's reveal post ever went up.

The reveal itself is only a partial confirmation. Z.ai's blog states plainly that "we tested GLM-5.3-Flash anonymously as ox-alpha... to gather user feedback" — that settles product identity. It does not settle checkpoint-exact identity: LiveBench's "ox-alpha-max" row, scored 69.2 overall, remains unrenamed as of 2026-08-26, so this site does not attribute that score to GLM-5.3-Flash.

The model's license is in dispute the same day it launched. Z.ai lists GLM-5.3-Flash as open-weight under MIT, backed by a real, downloadable HuggingFace repository. Artificial Analysis's own page for the model disagrees — it describes GLM-5.3-Flash as proprietary, text-only, and capped at a 400K context window, versus the 1,048,576-token window Zhipu claims. This site sides with the HuggingFace evidence, since a repo a reader can download and run is hard to argue with, but the disagreement is worth disclosing rather than quietly picking a side.

Two prices and a download link

Only two of the three models have a hosted API price to compare at all.

ModelList price (in / out, per 1M tokens)
GLM-5.3-Flash$0.15 / $0.50 — promotional $0.075 / $0.25 through 2026-09-09 24:00 UTC+8
DeepSeek V4 Flash$0.44 / $1.32 peak (01:00–04:00 and 06:00–10:00 UTC Mon–Fri — roughly 09:00–12:00 and 14:00–18:00 Beijing time); $0.22 / $0.66 all other hours, weekends included
Qwen3.8-Flash-Nextno hosted price exists

DeepSeek's pricing isn't a launch-day number — it took effect 2026-08-16, a genuine increase from a flatter ~$0.14/$0.28 rate at launch. Even at its cheapest off-peak window, $0.22/$0.66 sits above GLM-5.3-Flash's promotional $0.075/$0.25, and above GLM-5.3-Flash's post-promotion list price too.

Qwen3.8-Flash-Next isn't in that comparison because there's nothing to compare — it's download-and-self-host only, with no hosted endpoint anywhere. Worth flagging separately: a $0.16/$0.47 price does circulate online attached to the name "Qwen3.8-Flash." That's a different, not-yet-released hosted product Alibaba's own model card describes in future tense — not the open-weight preview covered here, and not a price this article uses for it.

That's a real three-way difference in go-to-market strategy, not a hole in this site's data: Alibaba shipped weights under the Qwen Community License 1.0, a custom permissive license rather than the OSI Apache-2.0 template, headlined at 125B total / 6B active parameters, though the full checkpoint on disk runs closer to 180B once its 51B-parameter n-gram embedding and 4B-parameter multi-token-prediction head are counted. The architecture underneath is itself labeled a Qwen4-generation preview — the model card names its class Qwen4ExpForConditionalGeneration — shipped under a "Qwen3.8" version number. Its context window, 262,144 tokens natively, is also the smallest of the three tracked here; the card claims extensibility to 1M, but that isn't the figure this site records. Running any of it is on the reader, not a hosted API.

What the provenance gap tells you

Not a winner. This site doesn't publish a composite score across benchmarks, and nothing here changes that. What the numbers do support: DeepSeek V4 Flash has a fully independent record on every benchmark tracked for it, at the highest hosted price of the three — more than GLM-5.3-Flash even at GLM's post-promotion list rate, let alone its current promo. GLM-5.3-Flash sits in between on verification: two-thirds of its tracked scores are independently sourced, against zero for Qwen3.8-Flash-Next, and it arrives with a live licensing dispute plus a six-day stealth run worth watching. Qwen3.8-Flash-Next has nothing an outside evaluator has confirmed yet — that could change the moment one of the eight boards this site checks picks it up. Full records, including every score's date and source, live on each model's own page and update there first: GLM-5.3-Flash, Qwen3.8-Flash-Next, DeepSeek V4 Flash.


Sources: benchmark scores and provenance tags from this site's own tracked data for GLM-5.3-Flash, Qwen3.8-Flash-Next, and DeepSeek V4 Flash. Ox Alpha reveal quote and licensing claim from Z.ai's own blog. Conflicting license description from Artificial Analysis's GLM-5.3-Flash page. DeepSWE v1.1 scores cross-checked against the deepswe.datacurve.ai official leaderboard. Qwen3.8-Flash-Next's benchmark figures and parameter breakdown come from its HuggingFace model card; its absence from independent boards checked across Artificial Analysis, vals.ai, ARC Prize, LiveBench, MathArena, DeepSWE's own leaderboard, Toolathlon, and Snorkel AI as of 2026-08-26.