Alibaba

Qwen3.8-Max

Eight of the ten benchmark scores tracked here, including the core coding results (SWE-bench Verified, LiveCodeBench, DeepSWE, Terminal-Bench 2.1), are independently confirmed — only the Toolathlon-Verified number behind Alibaba's agentic claim, and HLE with tools, are still self-reported. The hosted endpoint has served the Qwen3.8-Max-0902 snapshot since 2026-09-05; rows observed before that were earned by the 0803 build.

Qwen3.8-Max benchmarks and pricing, every number sourced: 10 tracked Qwen3.8-Max benchmark scores (8 independently run, 2 still resting on a vendor’s own claim), priced at $2.00 per million input tokens and $6.00 per million output.

Qwen3.8-Max’s 10 benchmark scores on this page were each verified against their sources between 2026-08-13 and 2026-09-29.

Released
2026-08-03
License
proprietary
Context window
1M tokens
Knowledge cutoff
Not disclosed
Parameters
Not disclosed
Architecture
Not disclosed

Qwen3.8-Max’s verified record

Qwen3.8-Max’s most-compared rival is Kimi K3: 2 trails, 2 ties, and 6 not callable across their 10 shared comparisons. Qwen3.8-Max is priced at $2.00/$6.00 per 1M tokens (in/out) vs Kimi K3’s $3.00/$15.00.

Against the 219 head-to-head comparisons Qwen3.8-Max shares with other tracked models: 55 real gaps, 42 inside the noise band, and 122 we will not call.

A gap counts for Qwen3.8-Max only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Qwen3.8-Max trails on 39 of them.

HLE · no tools Reasoning

±2 is noise
Behind
Step 5 Preview −3.4 · Kimi K3 −3.8 · Gemini 3.1 Pro Preview −3.9 · Gemini 3.8 Flash −4.7 · GPT-6 Sol −4.8 · Muse Spark 1.3 −5.6 · MiMo-V2.6-Pro −6.3 · GPT-5.6 Sol −6.4 · GPT-6.1 Sol −9.8 · GPT-6 Astra −11.6 · Claude Opus 5 −11.8 · Claude Sonnet 5.5 −11.9 · Claude Fable 5.1 −16.0 · Claude Opus 5.5 −18.3 · 4 superseded: Muse Spark 1.2 −2.4 · Gemini 3.7 Flash −4.8 · Claude Opus 4.8 −5.6 · Claude Fable 5 −12.4
Ahead
DeepSeek V4 Pro (0813) +2.1 · GLM-5.3-Flash +3.2 · GPT-5.6 Luna +3.6 · GPT-6 Luna +4.6 · Qwen3.8-Flash-Next +5.1 · 2 superseded: GLM-5.2 +2.0 · DeepSeek V4 Flash (0731) +4.5
Tie
4 models within ±2
Unverified
2 models — vendor-reported on one side

LiveBench Composite score across 7 domains

±2.7 is noise
Behind
Muse Spark 1.3 −3.1 · GPT-6.1 Sol −3.1 · Claude Opus 5.5 −4.7 · Claude Fable 5.1 −4.9 · 1 superseded: Claude Fable 5 −4.5
Ahead
GPT-5.6 Luna +4.9 · GPT-6 Luna +6.5 · GLM-5.3-Flash +6.9 · MiniMax M3 +11.2 · 2 superseded: DeepSeek V4 Flash (0731) +4.3 · GLM-5.2 +5.3
Tie
15 models within ±2.7

DeepSWE Long-horizon coding

±9.5 is noise
Behind
Grok 4.6 −10.0 · GPT-5.6 Luna −10.0 · Kimi K3 −12.0 · GLM-5.3 −12.0 · GPT-5.6 Sol −16.0 · Claude Opus 5 −17.0 · Gemini 3.8 Flash −17.0 · GPT-6 Astra −17.0 · 1 superseded: Claude Fable 5 −13.0
Ahead
1 superseded: GLM-5.2 +13.0
Tie
6 models within ±9.5
Unverified
10 models — vendor-reported on one side
Setup-dependent
2 models — scored on a different harness

Agents' Last Exam Professional work

±3.2 is noise
Behind
GPT-5.6 Luna −3.3 · GPT-5.6 Sol −3.6 · Claude Opus 5 −5.2 · GPT-6 Sol −5.2 · Muse Spark 1.3 −5.2 · GPT-6 Astra −7.2 · Claude Opus 5.5 −11.2
Ahead
Gemini 3.1 Pro Preview +10.6 · 1 superseded: GLM-5.2 +6.6
Tie
4 models within ±3.2
Unverified
6 models — vendor-reported on one side

No verdict for Qwen3.8-Max anywhere on Terminal-Bench 2.1 (every independently confirmed comparison inside the noise band); GPQA Diamond, LiveCodeBench, SWE-bench Verified (saturated); HLE · with tools, Toolathlon-Verified (nothing independently confirmed on both sides).

Qwen3.8-Max API pricing

$2.00 in / $6.00 out per 1M tokens — official pricing

What Qwen3.8-Max costs per job

Qwen3.8-Max cost for three reference workloads, computed from its list rates
WorkloadTokens in / outCost
One long chat turn100K / 10K$0.260
A codebase review1,000K / 100K$2.60
A day of agent work10,000K / 1,000K$26.00

Computed from Qwen3.8-Max’s list rates above — cache discounts and batch tiers are not applied.

Qwen3.8-Max is one of 2 Alibaba models tracked on this site, at these official list prices.

Alibaba model pricing, official list rates
ModelIn / 1MOut / 1M
Qwen3.8-Flash-Next——
Qwen3.8-Max$2.00$6.00

Qwen3.8-Max benchmark scores

Qwen3.8-Max benchmark scores, provenance, and source links
BenchmarkScore
HLE(no tools)[1]
Reasoning · ±2 is noise
HLE(with tools)
Reasoning · ±2 is noise
GPQA Diamondsaturated[2]
Expert science Q&A — not ranked at any gap size
Terminal-Bench 2.1[3]
Terminal ops · ±10.6 is noise
DeepSWE[4]
Long-horizon coding · ±9.5 is noise
Toolathlon-Verified[5]
Multi-tool chores · ±9.7 is noise
Agents' Last Exam[6]
Professional work · ±3.2 is noise
SWE-bench Verifiedsaturated[7]
Bug fixing — not ranked at any gap size
LiveCodeBenchsaturated[8]
Contest coding — not ranked at any gap size
LiveBench[9]
Composite score across 7 domains · ±2.7 is noise

Who ran these numbers: 8 of 10 independent — vals.ai (3), artificialanalysis.ai (2), deepswe.datacurve.ai (1), snorkel.ai (1), livebench.ai (1); vendor self-reported (2).

  1. HLE: AA's own run of 'Qwen3.8 Max (0902)' (text-only, no tools; underlying 0.4310), the snapshot the hosted endpoint has served since 2026-09-05. AA's separate entry for the 0803 build (slug qwen3-8-max-0803) read 43.0, and Alibaba's self-reported launch figure for that build was 43.6 — all three agree closely.
  2. GPQA Diamond: vals.ai run, rank 5/133 (updated 2026-08-15). Replaces Alibaba's self-reported 92.6.
  3. Terminal-Bench 2.1: AA's own run of 'Qwen3.8 Max (0902)' at max effort (underlying 0.8876; Terminus 2 harness). Replaces Alibaba's self-reported 86.6 for the 0803 build (Claude Code harness, avg@10); AA's entry for the 0803 build reads 81.3. Not on tbench.ai's official board as of 2026-09-29. Cross-check (2026-10-01): vals.ai's archived Terminus 2 table lists 'Qwen 3.8 Max' at 67.42 without naming a build.
  4. DeepSWE: Independent score (57% ± 3%, $3.73/task) — debuted as the highest-scoring new entry on this leaderboard at the time.
  5. Toolathlon-Verified: Vendor-reported. Does not appear on the official Toolathlon-Verified leaderboard as of Aug 2026.
  6. Agents' Last Exam: Overall pass rate 27.0 (Claude Code, XHigh; score 52.5; $486). Corrects an earlier version of this page, which showed 52.5 — that was the partial-credit score, not the pass rate.
  7. SWE-bench Verified: vals.ai run, rank 13/83, bash-only harness; slowest of the top group at 41m42s/test (updated 2026-08-14).
  8. LiveCodeBench: vals.ai run, rank 8/138 (updated 2026-08-15).
  9. LiveBench: Board row "Qwen 3.8 Max" on the 2026-06-25 LiveBench release.

Notes on the record

The $2/$6 per-million-token price shown on this page is Alibaba Cloud's Singapore/international endpoint rate. China (Beijing) region pricing is lower ($1.65/$4.951 per 1M tokens) — and per Alibaba Cloud's pricing page (checked 2026-08-20), so is pricing on the Frankfurt (Germany), Virginia (US), Tokyo (Japan), and Hong Kong endpoints; only the Singapore endpoint bills at the higher $2/$6 rate, so "international" pricing is really a Singapore-specific surcharge rather than a blanket non-China rate.

Knowledge cutoff not officially disclosed by Alibaba. The proprietary hosted API tracked on this page is a separate artifact from Alibaba's open-weight release of the same underlying mixture-of-experts architecture (2.4 trillion parameters total, 95 billion active per token), published on Hugging Face as Qwen/Qwen3.8-2.4T-A95B under a custom "qwen3.8-max" license rather than Apache 2.0; that open release is text-only (no vision/video input), requires thinking mode for all interactions, and has a smaller native context window (262,144 tokens, extensible to ~1,010,000) than the hosted API's 1,000,000-token default (source: huggingface.co/Qwen/Qwen3.8-2.4T-A95B).

Snapshot (checked 2026-09-29): Alibaba released Qwen3.8-Max-0902 on 2026-09-02 — a post-training upgrade with the architecture, context window and price unchanged — and Model Studio's update notice moved the qwen3.8-max endpoint to that snapshot on 2026-09-05 10:00 UTC+8 with billing unchanged. This page tracks the endpoint, so it now describes the 0902 build. Artificial Analysis scores the two builds separately: HLE 43.0 → 43.1, GPQA Diamond 92.7 → 92.8, Terminal-Bench 2.1 81.3 → 88.8 (max effort). The HLE and Terminal-Bench 2.1 rows below are AA's 0902 runs; the vals.ai, DeepSWE, Agents' Last Exam and LiveBench rows were observed before the switch and belong to the 0803 build unless the source names the snapshot.

Compare with

FAQ

Has Qwen3.8-Max been independently benchmarked?

Of the ten scores tracked for Qwen3.8-Max on this site, eight come from independent evaluators — Artificial Analysis (HLE no tools and Terminal-Bench 2.1, both on the 0902 snapshot), Vals.ai (GPQA Diamond, SWE-bench Verified, LiveCodeBench), the DeepSWE leaderboard, Snorkel AI's Agents' Last Exam, and LiveBench's own board — while two are self-reported by Alibaba, sourced to the Qwen3.8-2.4T-A95B model card on Hugging Face: HLE with tools and Toolathlon-Verified. The core software-engineering benchmarks (SWE-bench Verified, LiveCodeBench, DeepSWE) all fall on the independent side, and Terminal-Bench 2.1 joined them when Artificial Analysis ran the 0902 build; Toolathlon-Verified is the one agentic number still vendor-reported.

Is Qwen3.8-Max open source or a proprietary model?

The Qwen3.8-Max tracked on this page is the pay-per-token model on Alibaba Cloud Model Studio — proprietary, not open weight. Alibaba separately published the underlying architecture on Hugging Face as Qwen/Qwen3.8-2.4T-A95B, but under a custom "qwen3.8-max" license rather than Apache 2.0 (unlike the smaller Qwen3.8-27B, which is Apache 2.0 — source: huggingface.co/Qwen/Qwen3.8-27B), and that release drops vision/video input, forces thinking mode on, and caps native context at 262,144 tokens versus the hosted API's 1,000,000-token default. Source: huggingface.co/Qwen/Qwen3.8-2.4T-A95B.

Does Qwen3.8-Max cost the same price in every region?

No. The $2 in / $6 out per-million-token price on this page is Alibaba Cloud's Singapore endpoint rate. Its Beijing endpoint bills $1.65 in / $4.951 out, and — per the same Alibaba Cloud pricing page, checked 2026-08-20 — the Frankfurt, Virginia (US), Tokyo, and Hong Kong endpoints bill at that same lower rate. Singapore is the outlier, not the norm: of Qwen3.8-Max's six live regions, it's the only one billing at the higher rate.

What is Qwen3.8-Max's knowledge cutoff date?

Alibaba has not published one. Neither the Model Studio documentation for the hosted Qwen3.8-Max API nor the Hugging Face card for the related open-weight Qwen3.8-2.4T-A95B checkpoint states a training-data cutoff, even though both list parameter counts, context limits, and benchmark tables in detail.

Why is Qwen3.8-Max also listed as Qwen3.8-2.4T-A95B on Hugging Face?

"2.4T-A95B" is Alibaba's shorthand for the mixture-of-experts design behind Qwen3.8-Max: 2.4 trillion parameters total, with 95 billion active per token. It's the repository name for the open-weight release of that architecture; the hosted, billed API tracked on this page is versioned separately as "Qwen3.8-Max." Source: huggingface.co/Qwen/Qwen3.8-2.4T-A95B.

Further reading

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.