Alibaba
Qwen3.8-Max
Eight of the ten benchmark scores tracked here, including the core coding results (SWE-bench Verified, LiveCodeBench, DeepSWE, Terminal-Bench 2.1), are independently confirmed — only the Toolathlon-Verified number behind Alibaba's agentic claim, and HLE with tools, are still self-reported. The hosted endpoint has served the Qwen3.8-Max-0902 snapshot since 2026-09-05; rows observed before that were earned by the 0803 build.
Qwen3.8-Max benchmarks and pricing, every number sourced: 10 tracked Qwen3.8-Max benchmark scores (8 independently run, 2 still resting on a vendor’s own claim), priced at $2.00 per million input tokens and $6.00 per million output.
Qwen3.8-Max’s 10 benchmark scores on this page were each verified against their sources between 2026-08-13 and 2026-09-29.
- Released
- 2026-08-03
- License
- proprietary
- Context window
- 1M tokens
- Knowledge cutoff
- Not disclosed
- Parameters
- Not disclosed
- Architecture
- Not disclosed
Qwen3.8-Max’s verified record
Qwen3.8-Max’s most-compared rival is Kimi K3: 2 trails, 2 ties, and 6 not callable across their 10 shared comparisons. Qwen3.8-Max is priced at $2.00/$6.00 per 1M tokens (in/out) vs Kimi K3’s $3.00/$15.00.
Against the 219 head-to-head comparisons Qwen3.8-Max shares with other tracked models: 55 real gaps, 42 inside the noise band, and 122 we will not call.
A gap counts for Qwen3.8-Max only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Qwen3.8-Max trails on 39 of them.
HLE · no tools Reasoning
±2 is noiseLiveBench Composite score across 7 domains
±2.7 is noiseDeepSWE Long-horizon coding
±9.5 is noiseAgents' Last Exam Professional work
±3.2 is noiseNo verdict for Qwen3.8-Max anywhere on Terminal-Bench 2.1 (every independently confirmed comparison inside the noise band); GPQA Diamond, LiveCodeBench, SWE-bench Verified (saturated); HLE · with tools, Toolathlon-Verified (nothing independently confirmed on both sides).
Qwen3.8-Max API pricing
$2.00 in / $6.00 out per 1M tokens — official pricing
What Qwen3.8-Max costs per job
| Workload | Tokens in / out | Cost |
|---|---|---|
| One long chat turn | 100K / 10K | $0.260 |
| A codebase review | 1,000K / 100K | $2.60 |
| A day of agent work | 10,000K / 1,000K | $26.00 |
Computed from Qwen3.8-Max’s list rates above — cache discounts and batch tiers are not applied.
Qwen3.8-Max is one of 2 Alibaba models tracked on this site, at these official list prices.
| Model | In / 1M | Out / 1M |
|---|---|---|
| Qwen3.8-Flash-Next | — | — |
| Qwen3.8-Max | $2.00 | $6.00 |
Qwen3.8-Max benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise | |
| HLE(with tools) Reasoning · ±2 is noise | |
| GPQA Diamondsaturated[2] Expert science Q&A — not ranked at any gap size | |
| Terminal-Bench 2.1[3] Terminal ops · ±10.6 is noise | |
| DeepSWE[4] Long-horizon coding · ±9.5 is noise | |
| Toolathlon-Verified[5] Multi-tool chores · ±9.7 is noise | |
| Agents' Last Exam[6] Professional work · ±3.2 is noise | |
| SWE-bench Verifiedsaturated[7] Bug fixing — not ranked at any gap size | |
| LiveCodeBenchsaturated[8] Contest coding — not ranked at any gap size | |
| LiveBench[9] Composite score across 7 domains · ±2.7 is noise |
Who ran these numbers: 8 of 10 independent — vals.ai (3), artificialanalysis.ai (2), deepswe.datacurve.ai (1), snorkel.ai (1), livebench.ai (1); vendor self-reported (2).
- HLE: AA's own run of 'Qwen3.8 Max (0902)' (text-only, no tools; underlying 0.4310), the snapshot the hosted endpoint has served since 2026-09-05. AA's separate entry for the 0803 build (slug qwen3-8-max-0803) read 43.0, and Alibaba's self-reported launch figure for that build was 43.6 — all three agree closely.
- GPQA Diamond: vals.ai run, rank 5/133 (updated 2026-08-15). Replaces Alibaba's self-reported 92.6.
- Terminal-Bench 2.1: AA's own run of 'Qwen3.8 Max (0902)' at max effort (underlying 0.8876; Terminus 2 harness). Replaces Alibaba's self-reported 86.6 for the 0803 build (Claude Code harness, avg@10); AA's entry for the 0803 build reads 81.3. Not on tbench.ai's official board as of 2026-09-29. Cross-check (2026-10-01): vals.ai's archived Terminus 2 table lists 'Qwen 3.8 Max' at 67.42 without naming a build.
- DeepSWE: Independent score (57% ± 3%, $3.73/task) — debuted as the highest-scoring new entry on this leaderboard at the time.
- Toolathlon-Verified: Vendor-reported. Does not appear on the official Toolathlon-Verified leaderboard as of Aug 2026.
- Agents' Last Exam: Overall pass rate 27.0 (Claude Code, XHigh; score 52.5; $486). Corrects an earlier version of this page, which showed 52.5 — that was the partial-credit score, not the pass rate.
- SWE-bench Verified: vals.ai run, rank 13/83, bash-only harness; slowest of the top group at 41m42s/test (updated 2026-08-14).
- LiveCodeBench: vals.ai run, rank 8/138 (updated 2026-08-15).
- LiveBench: Board row "Qwen 3.8 Max" on the 2026-06-25 LiveBench release.
Notes on the record
The $2/$6 per-million-token price shown on this page is Alibaba Cloud's Singapore/international endpoint rate. China (Beijing) region pricing is lower ($1.65/$4.951 per 1M tokens) — and per Alibaba Cloud's pricing page (checked 2026-08-20), so is pricing on the Frankfurt (Germany), Virginia (US), Tokyo (Japan), and Hong Kong endpoints; only the Singapore endpoint bills at the higher $2/$6 rate, so "international" pricing is really a Singapore-specific surcharge rather than a blanket non-China rate.
Knowledge cutoff not officially disclosed by Alibaba. The proprietary hosted API tracked on this page is a separate artifact from Alibaba's open-weight release of the same underlying mixture-of-experts architecture (2.4 trillion parameters total, 95 billion active per token), published on Hugging Face as Qwen/Qwen3.8-2.4T-A95B under a custom "qwen3.8-max" license rather than Apache 2.0; that open release is text-only (no vision/video input), requires thinking mode for all interactions, and has a smaller native context window (262,144 tokens, extensible to ~1,010,000) than the hosted API's 1,000,000-token default (source: huggingface.co/Qwen/Qwen3.8-2.4T-A95B).
Snapshot (checked 2026-09-29): Alibaba released Qwen3.8-Max-0902 on 2026-09-02 — a post-training upgrade with the architecture, context window and price unchanged — and Model Studio's update notice moved the qwen3.8-max endpoint to that snapshot on 2026-09-05 10:00 UTC+8 with billing unchanged. This page tracks the endpoint, so it now describes the 0902 build. Artificial Analysis scores the two builds separately: HLE 43.0 → 43.1, GPQA Diamond 92.7 → 92.8, Terminal-Bench 2.1 81.3 → 88.8 (max effort). The HLE and Terminal-Bench 2.1 rows below are AA's 0902 runs; the vals.ai, DeepSWE, Agents' Last Exam and LiveBench rows were observed before the switch and belong to the 0803 build unless the source names the snapshot.
Compare with
FAQ
Has Qwen3.8-Max been independently benchmarked?
Of the ten scores tracked for Qwen3.8-Max on this site, eight come from independent evaluators — Artificial Analysis (HLE no tools and Terminal-Bench 2.1, both on the 0902 snapshot), Vals.ai (GPQA Diamond, SWE-bench Verified, LiveCodeBench), the DeepSWE leaderboard, Snorkel AI's Agents' Last Exam, and LiveBench's own board — while two are self-reported by Alibaba, sourced to the Qwen3.8-2.4T-A95B model card on Hugging Face: HLE with tools and Toolathlon-Verified. The core software-engineering benchmarks (SWE-bench Verified, LiveCodeBench, DeepSWE) all fall on the independent side, and Terminal-Bench 2.1 joined them when Artificial Analysis ran the 0902 build; Toolathlon-Verified is the one agentic number still vendor-reported.
Is Qwen3.8-Max open source or a proprietary model?
The Qwen3.8-Max tracked on this page is the pay-per-token model on Alibaba Cloud Model Studio — proprietary, not open weight. Alibaba separately published the underlying architecture on Hugging Face as Qwen/Qwen3.8-2.4T-A95B, but under a custom "qwen3.8-max" license rather than Apache 2.0 (unlike the smaller Qwen3.8-27B, which is Apache 2.0 — source: huggingface.co/Qwen/Qwen3.8-27B), and that release drops vision/video input, forces thinking mode on, and caps native context at 262,144 tokens versus the hosted API's 1,000,000-token default. Source: huggingface.co/Qwen/Qwen3.8-2.4T-A95B.
Does Qwen3.8-Max cost the same price in every region?
No. The $2 in / $6 out per-million-token price on this page is Alibaba Cloud's Singapore endpoint rate. Its Beijing endpoint bills $1.65 in / $4.951 out, and — per the same Alibaba Cloud pricing page, checked 2026-08-20 — the Frankfurt, Virginia (US), Tokyo, and Hong Kong endpoints bill at that same lower rate. Singapore is the outlier, not the norm: of Qwen3.8-Max's six live regions, it's the only one billing at the higher rate.
What is Qwen3.8-Max's knowledge cutoff date?
Alibaba has not published one. Neither the Model Studio documentation for the hosted Qwen3.8-Max API nor the Hugging Face card for the related open-weight Qwen3.8-2.4T-A95B checkpoint states a training-data cutoff, even though both list parameter counts, context limits, and benchmark tables in detail.
Why is Qwen3.8-Max also listed as Qwen3.8-2.4T-A95B on Hugging Face?
"2.4T-A95B" is Alibaba's shorthand for the mixture-of-experts design behind Qwen3.8-Max: 2.4 trillion parameters total, with 95 billion active per token. It's the repository name for the open-weight release of that architecture; the hosted, billed API tracked on this page is versioned separately as "Qwen3.8-Max." Source: huggingface.co/Qwen/Qwen3.8-2.4T-A95B.
Further reading
- GPQA Diamond leaderboard 2026 — Qwen3.8-Max is one of the 23 models it compares.
- Models with 10M token context windows 2026 — Qwen3.8-Max is one of the 34 models it compares.