Three cheap Chinese models, only one fully independent scorecard
August 27, 2026 · Updated October 6, 2026
GLM-5.3-Flash, Qwen3.8-Flash-Next, and DeepSeek V4 Flash are three open-weight (GLM-5.3-Flash's status is disputed — more below) budget-tier models from three different Chinese labs, all released within about a month of each other. Between them, this site tracks 26 benchmark line items. Every one of DeepSeek's ten comes from an independent evaluator. GLM-5.3-Flash's nine split eight independent and one self-reported. Qwen3.8-Flash-Next launched with four scores, all self-reported — on 2026-08-26 this site could not find it on a single one of eight independent leaderboards; within two days that had flipped to four independent of seven, the fastest provenance turnaround this site had recorded to that point. The scores below are real. Who ran them is the actual comparison.
The split that doesn't average out
Independent-score ratio as of 2026-08-28 (superseded by the October 6 counts above): DeepSeek V4 Flash 8 of 8 (100%), GLM-5.3-Flash 5 of 7 (71%), Qwen3.8-Flash-Next 4 of 7 (57%) — up from 0 of 4 at launch.
DeepSeek is the elder of the three by about four weeks — it shipped 2026-07-31, while GLM-5.3-Flash and Qwen3.8-Flash-Next both landed 2026-08-26. That head start doesn't explain the gap on its own: independent boards have had a full month to score DeepSeek, but they've also had roughly six days to score GLM-5.3-Flash under its former alias (more on that below), and — counting from when its downloadable weights actually appeared, not its official reveal date — roughly two days for Qwen3.8-Flash-Next, whose Hugging Face repo went up 2026-08-24.
Independent doesn't mean better. It means a party with no stake in the outcome ran the eval and published the number. A self-reported score can still be accurate — and the update above shows what happens when outside evaluators arrive: Artificial Analysis's independent HLE run came in at 38.0 against Alibaba's thinking-mode 35.9, a different harness producing a different number, which is exactly why the two were never averaged here. Three of Qwen's agentic claims — DeepSWE, Toolathlon-Verified, Agents' Last Exam — remain vendor-only as of 2026-08-28.
Six benchmarks, all three models
Six of the fourteen benchmarks this site tracks have a published number for all three models:
| Benchmark | DeepSeek V4 Flash | GLM-5.3-Flash | Qwen3.8-Flash-Next |
|---|---|---|---|
| HLE (no tools) | 38.6 — independent | 39.9 — independent | 38.0 — independent (AA, added 2026-08-28) |
| DeepSWE v1.1 | 53.0 — independent | 63% — independent (official board, "max" tier) | 58.7 — self-reported |
| Toolathlon-Verified | 70.7 — independent | 78.4 — independent (official board, 2026-10-03) | 73.5 — self-reported (Pass@1) |
| GPQA Diamond | 89.9 — independent | 91.2 — independent | 92.3 — independent (AA, added 2026-08-28) |
| LiveBench | 74.2 — independent | 71.6 — independent (added 2026-08-28) | 76.2 — independent (added 2026-08-28) |
| Terminal-Bench 2.1 | 78.65 — independent (AA, added 2026-10-03) | 84.3 — independent (AA, 2026-08-26) | 86.1 — independent (AA, added 2026-08-28) |
Two caveats before reading anything into those rows. (Qwen's original vendor HLE — 35.9, thinking mode, GPT-4o judge — was superseded on 2026-08-28 by Artificial Analysis's independent 38.0 no-tools run, so the HLE row is now a clean three-way independent match.) GLM-5.3-Flash's DeepSWE figure is pulled from the official board's "max" tier; DeepSeek's row doesn't specify a tier, so a 10-point gap between two independently-sourced numbers may not be as clean as it looks. And only Qwen's Toolathlon-Verified score is explicitly labeled Pass@1 — DeepSeek's and GLM-5.3-Flash's aren't, so there's no confirmation all three Toolathlon scores measure the same pass criterion.
Narrow the table to rows where both DeepSeek and GLM-5.3-Flash carry an independent number — all six, once GLM-5.3-Flash's Toolathlon row was upgraded to the official board's own run on 2026-10-03 and its Terminal-Bench 2.1 row arrived with this table — and GPQA Diamond drops out first: this site grades it saturated, so per its own methodology that gap never drives a verdict regardless of who ran it (the full case for that downgrade is its own piece). Five rows remain, and still only one clears this site's noise band. On HLE, GLM-5.3-Flash's 39.9 against DeepSeek's 38.6 is a 1.3-point gap on a 2,500-question benchmark whose meaningful-gap threshold is 2.0 points — inside the band, so it logs as a tie, not a lead, no matter which name is on top. On DeepSWE v1.1 the 10-point gap does clear that benchmark's 9.5-point threshold, but only just, and the tier caveat above still applies — so treat it as provenance-clean, not necessarily measurement-clean. On LiveBench, DeepSeek's 74.2 against GLM-5.3-Flash's 71.6 is a 2.6-point gap against that board's 2.7-point threshold — inside the band by a tenth of a point, so also a tie. On Terminal-Bench 2.1, GLM-5.3-Flash's 84.3 (Artificial Analysis) against DeepSeek's 78.65 is a 5.65-point gap inside that benchmark's 10.6-point band — a tie again. Toolathlon-Verified now posts independent numbers on both sides — DeepSeek's 70.7 (vals.ai) against GLM-5.3-Flash's 78.4 (the official board's own run, recorded 2026-10-03) — and the 7.7-point spread sits inside that benchmark's 9.7-point band, so the row logs as a tie with GLM-5.3-Flash's number provenance-upgraded rather than the matchup settled. None of this is a verdict; it's just what's left once the self-reported numbers are set aside.
One benchmark sits outside the three-way table: GLM-5.3-Flash and Qwen3.8-Flash-Next both have Agents' Last Exam — 26.3 versus 24.3, both self-reported — while DeepSeek V4 Flash has no row there at all. The same pass-criterion gap applies here as above: Qwen's figure is labeled Pass@1, GLM-5.3-Flash's isn't, so it's not confirmed both numbers measure the same thing. With provenance on neither side either way, that pair logs as unverified rather than a real gap or a tie.
Ox Alpha, unmasked: GLM-5.3-Flash's stealth run
GLM-5.3-Flash's launch carries a second story: it's the model that spent roughly six days running anonymously on OpenRouter and OpenCode as "ox-alpha," starting around 2026-08-20. This site's own search data showed real reader demand for "ox alpha" benchmark queries before Zhipu's reveal post ever went up.
The reveal itself is only a partial confirmation. Z.ai's blog states plainly that "we tested GLM-5.3-Flash anonymously as ox-alpha... to gather user feedback" — that settles product identity. It does not settle checkpoint-exact identity — and LiveBench went on to prove the distinction matters. On 2026-08-28 it added a separate "GLM-5.3 Flash" row at 71.6 overall while keeping the stealth-period "ox-alpha-max" row (69.2) unrenamed beside it: the GA build and the preview build score 2.4 points apart on the same board. This site never attributed the 69.2 to GLM-5.3-Flash; the 71.6 is now on its page.
The model's license is in dispute the same day it launched. Z.ai lists GLM-5.3-Flash as open-weight under MIT, backed by a real, downloadable HuggingFace repository. Artificial Analysis's own page for the model disagrees — it describes GLM-5.3-Flash as proprietary, text-only, and capped at a 400K context window, versus the 1,048,576-token window Zhipu claims. This site sides with the HuggingFace evidence, since a repo a reader can download and run is hard to argue with, but the disagreement is worth disclosing rather than quietly picking a side.
GLM-5.3-Flash vs DeepSeek pricing — and Qwen's download link
Only two of the three models have a hosted API price to compare at all.
| Model | List price (in / out, per 1M tokens) |
|---|---|
| GLM-5.3-Flash | $0.15 / $0.50 — launch promo of $0.075 / $0.25 expired 2026-09-09 24:00 UTC+8 |
| DeepSeek V4 Flash | $0.44 / $1.32 peak (01:00–04:00 and 06:00–10:00 UTC Mon–Fri — roughly 09:00–12:00 and 14:00–18:00 Beijing time); $0.22 / $0.66 all other hours, weekends included |
| Qwen3.8-Flash-Next | no hosted price exists |
DeepSeek's pricing isn't a launch-day number — it took effect 2026-08-16, a genuine increase from a flatter ~$0.14/$0.28 rate at launch. Even at its cheapest off-peak window, $0.22/$0.66 sits above GLM-5.3-Flash's expired promotional rate of $0.075/$0.25, and above its $0.15/$0.50 list price too.
Qwen3.8-Flash-Next isn't in that comparison because there's nothing to compare — it's download-and-self-host only, with no hosted endpoint anywhere. Worth flagging separately: a $0.16/$0.47 price does circulate online attached to the name "Qwen3.8-Flash." That's a different hosted product — it went on sale through Alibaba's cloud API on 2026-08-26 — not the open-weight preview covered here, and not a price this article uses for it.
That's a real three-way difference in go-to-market strategy, not a hole in this site's data: Alibaba shipped weights under the Qwen Community License 1.0, a custom permissive license rather than the OSI Apache-2.0 template, headlined at 125B total / 6B active parameters, though the full checkpoint on disk runs closer to 180B once its 51B-parameter n-gram embedding and 4B-parameter multi-token-prediction head are counted. The architecture underneath is itself labeled a Qwen4-generation preview — the model card names its class Qwen4ExpForConditionalGeneration — shipped under a "Qwen3.8" version number. Its context window, 262,144 tokens natively, is also the smallest of the three tracked here; the card claims extensibility to 1M, but that isn't the figure this site records. Running any of it is on the reader, not a hosted API.
What the provenance gap tells you
Not a winner. This site doesn't publish a composite score across benchmarks, and nothing here changes that. What the numbers do support: DeepSeek V4 Flash has a fully independent record on every benchmark tracked for it, at the highest hosted price of the three — more than GLM-5.3-Flash even at its list rate of $0.15 / $0.50, let alone the $0.075 / $0.25 promo it charged through September 9. GLM-5.3-Flash sits in between on verification: eight of its nine tracked scores are independently sourced, and it arrives with a live licensing dispute plus a stealth run whose checkpoint question LiveBench has now answered. Qwen3.8-Flash-Next is the moving story — zero independent scores at launch, four of seven two days later, with only its agentic claims still resting on Alibaba's word. "That could change the moment one of the eight boards picks it up" is how this piece originally ended that sentence; it took 48 hours. Full records, including every score's date and source, live on each model's own page and update there first: GLM-5.3-Flash, Qwen3.8-Flash-Next, DeepSeek V4 Flash.
Common questions about the three
Which of the three models has the best-verified record?
DeepSeek V4 Flash (0731): every one of its ten tracked scores came from an independent evaluator — the only one of the three with no vendor-only rows left. The latest arrival was vals.ai's Terminal-Bench 4.0 run, recorded 2026-10-08.
Was GLM-5.3-Flash's license disputed?
Yes. Z.ai lists GLM-5.3-Flash as open-weight under MIT, backed by a downloadable HuggingFace repository, while Artificial Analysis's own page for the model describes it as proprietary, text-only, and capped at a 400K window. This piece discloses the disagreement rather than quietly picking a side.
Did Qwen3.8-Flash-Next ever get independent scores?
Yes, and quickly. At launch (2026-08-26) it appeared on none of the eight independent boards this site checked; within two days four of its seven tracked scores were independent runs — the fastest provenance turnaround this site had recorded to that point.
What was the Ox Alpha stealth run?
GLM-5.3-Flash ran on LiveBench unnamed, as the row this site tracked under ox-alpha-max. When the generally-available build arrived under its own name, the two sat 2.4 points apart on the same board — settling the stealth-checkpoint question this piece had flagged.
Which of the three is cheapest to run?
GLM-5.3-Flash, at its $0.15/$0.50 list rate per million tokens (a $0.075/$0.25 promo ran through September 9). DeepSeek V4 Flash is the dearest of the three at $0.44/$1.32 peak — and it is also the one with the fully independent record, so the price gap and the trust gap point in opposite directions.
Why does this comparison refuse to pick a winner?
Because this site publishes no composite score across benchmarks, and the honest summary is a split: DeepSeek wins on verification, GLM on price, Qwen on momentum. Who ran the numbers is the actual comparison.
Are these three models still current?
DeepSeek V4 Flash was superseded by V4.1 Flash per DeepSeek's own changelog on September 10 — its page stays live with the record it earned while current. GLM-5.3-Flash and Qwen3.8-Flash-Next have no computed successors, so their records above remain the live ones for their lines.
How many benchmark scores does this comparison cover?
26 tracked line items across the three models: DeepSeek V4 Flash ten, GLM-5.3-Flash nine, Qwen3.8-Flash-Next seven — each with its evaluator, date, and harness named in the tables above.
Revision log
- Updated October 8, 2026: two new tracked benchmarks added to the site (Terminal-Bench 4.0 and OSWorld 2.0) ripple through every count here — DeepSeek V4 Flash to ten tracked scores (all independent; the new one is vals.ai's Terminal-Bench 4.0 run), GLM-5.3-Flash to nine (eight independent), totals from 24 to 26 line items — and the three-way table gains Terminal-Bench 2.1, where all three carry independent numbers, as a sixth shared benchmark.
- Updated October 6, 2026: counts refreshed against the live data files — DeepSeek V4 Flash moves from eight tracked scores to nine (Terminal-Bench 2.1 78.65, an Artificial Analysis run recorded 2026-10-03), and GLM-5.3-Flash from seven to eight with its Toolathlon-Verified row upgraded to the official board's independent entry, so the provenance split now reads seven independent of eight rather than five of seven. The totals move from 22 to 24 line items; the provenance argument — who ran the numbers is the comparison — is unchanged. An FAQ section was added the same day.
- Update, August 28, 2026: the boards moved within two days of publication, exactly as this piece said they could. Artificial Analysis published a full Qwen3.8-Flash-Next page (HLE 38.0, GPQA Diamond 92.3, Terminal-Bench 2.1 86.1) and LiveBench added both models mid-cycle — Qwen3.8-Flash-Next at 76.2 and, decisively, a separate GLM-5.3 Flash row at 71.6 that sits beside the still-unrenamed ox-alpha-max row at 69.2. The ratios are now 5-of-7, 4-of-7, and 8-of-8, and the stealth-checkpoint question this piece flagged is settled: the GA build and the preview build score 2.4 points apart on the same board. The original argument — who ran the numbers is the comparison — is unchanged; the counts in the body above are updated in place.
- Update, September 10, 2026: DeepSeek marked V4 Flash superseded the same day it shipped V4.1 Flash, per DeepSeek's own changelog. That doesn't change the provenance analysis in the body above — the 8-of-8 independent-score record is what V4 Flash actually earned while it was current, and its page stays live with those scores intact. It does change the pricing table in the body above: DeepSeek is temporarily routing the
deepseek-v4-flashAPI name to V4.1 Flash at V4.1 Flash's lower rate, so a call using that name today bills differently than the $0.44/$1.32 peak rate quoted in that table, which is V4 Flash's own historical price. V4.1 Flash has its own page, and as of this update carries zero independent scores of its own. - Update, September 29, 2026: this site added a twelfth tracked benchmark, AA-AnalystAgent. None of the three models has a row on it (Artificial Analysis's board lists an earlier DeepSeek V4 Flash 0420 build, not the 0731 build tracked here), so the "five of" count in the body above now reads five of twelve and no other figure changed.
Sources: benchmark scores and provenance tags from this site's own tracked data for GLM-5.3-Flash, Qwen3.8-Flash-Next, and DeepSeek V4 Flash. Ox Alpha reveal quote and licensing claim from Z.ai's own blog. Conflicting license description from Artificial Analysis's GLM-5.3-Flash page. DeepSWE v1.1 scores cross-checked against the deepswe.datacurve.ai official leaderboard. Qwen3.8-Flash-Next's benchmark figures and parameter breakdown come from its HuggingFace model card; its absence from independent boards checked across Artificial Analysis, vals.ai, ARC Prize, LiveBench, MathArena, DeepSWE's own leaderboard, Toolathlon, and Snorkel AI as of 2026-08-26.