Alibaba
Qwen3.8-Flash-Next
Four days after shipping with nothing but Alibaba's own numbers, Qwen3.8-Flash-Next now has four independent scores out of seven tracked — Artificial Analysis published a full page (HLE 38.0, GPQA Diamond 92.3, Terminal-Bench 2.1 86.1) and LiveBench added it mid-cycle at 76.2 overall between 2026-08-26 and 2026-08-28. The three agentic headline claims, including the 73.5 Toolathlon-Verified and 58.7 DeepSWE figures, still rest on Alibaba's self-reports alone.
Qwen3.8-Flash-Next benchmarks and pricing, every number sourced: 7 tracked Qwen3.8-Flash-Next benchmark scores (4 independently run, 3 still resting on a vendor’s own claim), with no hosted API price — Qwen3.8-Flash-Next is open weights, self-host only.
Qwen3.8-Flash-Next architecture: Sparse Mixture-of-Experts; 125B MoE core total parameters (6B activated per token); 262K-token context window.
- Released
- 2026-08-26
- License
- open-weights
- Context window
- 262K tokens
- Knowledge cutoff
- Not disclosed
- Verified
- sources checked 2026-08-26–2026-08-28
- Parameters
- 125B MoE core (6B active)
- Architecture
- Sparse Mixture-of-Experts
Qwen3.8-Flash-Next’s verified record
Qwen3.8-Flash-Next’s most-compared rival is Kimi K3: 2 trails and 5 not callable across their 7 shared comparisons.
Against the 184 head-to-head comparisons Qwen3.8-Flash-Next shares with other tracked models: 42 real gaps, 34 inside the noise band, and 108 we will not call.
A gap counts for Qwen3.8-Flash-Next only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Qwen3.8-Flash-Next trails on 37 of them.
HLE · no tools Reasoning
±2 is noiseLiveBench Composite score across 7 domains
±2.7 is noiseNo verdict for Qwen3.8-Flash-Next anywhere on Terminal-Bench 2.1 ( every independently confirmed comparison inside the noise band); Agents' Last Exam, DeepSWE, Toolathlon-Verified ( nothing independently confirmed on both sides); GPQA Diamond ( saturated).
- None of Qwen3.8-Flash-Next’s coding comparisons are independently confirmed on both sides yet.
- No real gap yet in any of Qwen3.8-Flash-Next’s agentic comparisons — the independently confirmed ones all sit inside the noise band.
Qwen3.8-Flash-Next API pricing
— in / — out per 1M tokens — official pricing source
Qwen3.8-Flash-Next is one of 2 Alibaba models tracked on this site, at these official list prices.
| Model | In / 1M | Out / 1M |
|---|---|---|
| Qwen3.8-Flash-Next | — | — |
| Qwen3.8-Max | $2.00 | $6.00 |
Qwen3.8-Flash-Next benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise | |
| DeepSWE[2] Long-horizon coding · ±9.5 is noise | |
| Toolathlon-Verified[3] Multi-tool chores · ±9.7 is noise | |
| Agents' Last Exam[4] Professional work · ±3.2 is noise | |
| GPQA Diamondsaturated[5] Expert science Q&A — not ranked at any gap size | |
| Terminal-Bench 2.1[6] Terminal ops · ±10.6 is noise | |
| LiveBench[7] Composite score across 7 domains · ±2.7 is noise |
Who ran these numbers: 4 of 7 independent — artificialanalysis.ai (3), livebench.ai (1); vendor self-reported (3).
- HLE: Standard HLE (no tools), independently measured by Artificial Analysis, which published its Qwen3.8-Flash-Next page between this site's 2026-08-26 check (nothing listed) and 2026-08-28. Exact value from the page's underlying data. Replaces the vendor's self-reported 35.9, which was run in "thinking mode" and judged by GPT-4o per the HF card's own footnote — a different judge than AA uses, so the two figures were never directly comparable; the independent run supersedes it under this site's source-priority rule.
- DeepSWE: Card reports the higher of two harnesses (Claude Code vs. mini-SWE-agent; temp=1.0, top_p=0.95, 256K ctx) and states the model scores best on mini-SWE-agent -- the same harness the official board uses. Checked deepswe.datacurve.ai's v1.1 leaderboard live (last updated today, 2026-08-26): Qwen3.8-Flash-Next is not listed yet, only qwen3.8-max (57%±3%) appears for Alibaba. If later added, 58.7 would slot between claude-opus-4.8[max] (59%±2%) and qwen3.8-max[xhigh] (57%±3%) on that board.
- Toolathlon-Verified: Value is Pass@1. Vendor-reported Pass@1. Checked toolathlon.xyz's live Model Leaderboard (Toolathlon-Verified series): Qwen3.8-Flash-Next is not listed as of today; the independent top score there is Kimi K3 (max) at 76.5±1.9, and the only Qwen entry present is Qwen3.5 397B-A17B (40.7±2.0, dated 2026-07-30).
- Agents' Last Exam: Value is Pass@1 (a separate composite "Score" of 51.2 is also reported by the vendor but not used here). Vendor card reports both Pass@1 (24.3) and a separate composite 'Score' (51.2); Pass@1 used here since snorkel.ai's live leaderboard ranks primarily by Pass Rate. Checked snorkel.ai/leaderboard/agents-last-exam/ live: Qwen3.8-Flash-Next does not appear in the model filter dropdown as of today (only Qwen3 6-Plus, Qwen3 7-Max, Qwen3 8-27B, Qwen3 8-Max are listed for Alibaba).
- GPQA Diamond: Independently measured by Artificial Analysis (underlying value 0.923232...). First independent score for this model on any board this site tracks — AA's page appeared between the 2026-08-26 absence check and 2026-08-28. GPQA Diamond is graded saturated here, so this number never drives a verdict; it is recorded for the provenance ledger. AA's GPQA runs use its own harness, distinct from the vals.ai harness behind most of this column — compareScores() marks cross-harness pairs setup-dependent automatically.
- Terminal-Bench 2.1: Independently measured by Artificial Analysis (underlying value 0.861423...). The vendor's HF card reports no Terminal-Bench figure at all, so this is a wholly new data point rather than a confirmation of a self-report.
- LiveBench: LiveBench overall, read from the live board 2026-08-28 (row "Qwen 3.8 Flash Next", open, $0.042/run cost column). Notable on its own: this page previously said LiveBench could not add the model soon because its snapshot cadence is six months — LiveBench added it mid-cycle within days, so that structural claim was wrong and is corrected in the model notes.
Notes on the record
Qwen3.8-Flash-Next shipped as an open-weight-only architecture preview, not a production release (Hugging Face repo metadata: created 2026-08-24, last modified 2026-08-26). Despite the "Qwen3.8" version label, Hugging Face's own architecture metadata lists it as "Qwen4ExpForConditionalGeneration" (model_type "qwen4_exp") — the card frames it as an experimental preview of the architecture that will underpin Qwen4, not a Qwen3.8-generation model in the strict sense. No hosted API price exists for this exact checkpoint as of 2026-08-26 — no input or output price is listed anywhere, and a live check found no DashScope or Model Studio listing. The $0.16/1M input and $0.47/1M output figures circulating online belong to a different, not-yet-shipped sibling product: the bare "Qwen3.8-Flash" (without "-Next"), which the HF model card itself describes as "the official version based on Qwen3.8-Flash-Next with more production features, e.g., 1M context length by default, official built-in tools," now served via Alibaba's cloud API — it went on sale 2026-08-26 per OpenRouter's live listing of the maker-operated endpoint (the pricing itself was first posted by a Qwen/Alibaba_Qwen X post). This page's 262,144-token context window is the -Next preview's NATIVE length; the card separately says the preview is "extensible up to 1,000,000 tokens" via RoPE-scaling techniques such as YaRN, which is a different claim from the sibling's "1M context length by default" — extensible-with-work versus shipped-as-standard. Neither figure should be read as the other's. Licensing is "qwen-community-1.0" (Qwen Community License 1.0) per Hugging Face's structured API — a custom license granting broad commercial-use rights, distinct from the OSI-approved Apache-2.0 template; the same API confirms the checkpoint is gated:false, disabled:false, across 131 safetensors shards. Those shards also resolve an apparent size contradiction: Alibaba's headline "125B total, 6B active" describes only the Mixture-of-Experts core, while the full on-disk checkpoint totals roughly 180B parameters (179,999,981,459 across all dtypes — 179,999,981,424 of them BF16, per Hugging Face's safetensors metadata) once a 51B-parameter N-gram embedding layer and a 4B-parameter multi-token-prediction head are added in — both figures are accurate, they just answer different questions.
When this page first went up (2026-08-26), every score on it was self-reported by Alibaba and live checks found the model on none of eight independent boards. That changed within two days, faster than this page predicted: by 2026-08-28 Artificial Analysis had published a full Qwen3.8-Flash-Next page (HLE 38.0, GPQA Diamond 92.3, Terminal-Bench 2.1 86.1 — the first two replacing or sitting alongside vendor claims, the third a benchmark Alibaba never reported at all) and LiveBench had added it mid-cycle at 76.2 overall. Worth an explicit correction: this page previously argued LiveBench could not add the model soon because its snapshot cadence is six months — LiveBench added it within days anyway, so that structural claim was simply wrong. Three scores remain vendor-only as of 2026-08-28: DeepSWE 1.1 (58.7), Toolathlon Verified (73.5) and Agents' Last Exam (24.3) — toolathlon.xyz, deepswe.datacurve.ai and snorkel.ai still list no row for this model. The remaining vendor numbers carry methodology flags worth noting: the vendor's own 35.9 HLE (thinking mode, GPT-4o judge per the card's footnote) has now been superseded on this page by AA's independent 38.0 no-tools run — the two used different judges and were never directly comparable; the 58.7 DeepSWE 1.1 figure is the higher of two harnesses (Claude Code vs. mini-SWE-agent), and the card notes the model does best specifically on mini-SWE-agent — the same harness the official DeepSWE board itself uses, which is why a slot between claude-opus-4.8[max] (59%±2%) and qwen3.8-max[xhigh] (57%±3%) is a reasonable placement if the board, last updated the same day but still without this model, ever adds it; and the 24.3 Agents' Last Exam figure is Pass@1, chosen here over the vendor's separate 51.2 composite "Score" because Snorkel AI's own board ranks primarily by Pass Rate. The three still-vendor-only numbers should be read as claims pending independent re-confirmation on this exact checkpoint, not as settled scores.
Compare with
FAQ
What is Qwen3.8-Flash-Next actually good at, based on the available numbers?
The independent picture now covers general capability: Artificial Analysis measured HLE 38.0, GPQA Diamond 92.3 and Terminal-Bench 2.1 86.1, and LiveBench scores it 76.2 overall — mid-cycle additions published within two days of this page's launch check. The agentic claims are still Alibaba's own: 58.7 on DeepSWE 1.1 (the better of two test harnesses — mini-SWE-agent, the same harness the official DeepSWE board itself uses) and 73.5 Pass@1 on Toolathlon Verified, neither yet confirmed by the boards that run those benchmarks.
How much does Qwen3.8-Flash-Next cost to use via API?
The price you may have seen quoted — $0.16 per 1M input tokens, $0.47 per 1M output tokens — is not this model's price; that figure belongs to a separate sibling product, the bare "Qwen3.8-Flash" (without "-Next"), which went on sale through Alibaba's cloud API on 2026-08-26 with a 1M-token default window — OpenRouter lists the live maker-operated endpoint; we found no first-party Alibaba announcement page. It is a different product from the open-weight -Next preview this page tracks. As of 2026-08-26, Qwen3.8-Flash-Next itself has no DashScope or Model Studio listing at all.
Has Qwen3.8-Flash-Next been independently benchmarked yet?
Yes — and faster than this page originally predicted. At launch (2026-08-26) eight independent boards listed nothing; by 2026-08-28 Artificial Analysis had a full page (HLE 38.0, GPQA Diamond 92.3, Terminal-Bench 2.1 86.1) and LiveBench had added the model mid-cycle at 76.2 overall — despite this page having argued LiveBench's six-month snapshot cadence made that structurally unlikely, a claim events refuted within days. The agentic scores are the holdouts: DeepSWE, Toolathlon Verified and Agents' Last Exam still show only Alibaba's own figures as of 2026-08-28.
Is Qwen3.8-Flash-Next open source?
It is open-weights, released under a custom license called qwen-community-1.0 (Qwen Community License 1.0) — a permissive license granting broad commercial-use rights, distinct from the OSI-approved Apache-2.0 template. Hugging Face's structured metadata also lists the checkpoint as ungated and not disabled.
Why do different sources list this model's size as 6B, 125B, or 180B parameters?
None of the three numbers is wrong — each measures a different slice of the same checkpoint. The 6B and 125B figures both describe the Mixture-of-Experts core: 6B is what's active per token, 125B is that core's full size. The 180B figure is the entire on-disk checkpoint, which stacks two more components on top of that core — a large embedding layer used for cheap parameter scaling and a smaller head that predicts multiple tokens at once. Hugging Face's safetensors metadata puts the grand total at 179,999,981,459 parameters, all but 35 of them BF16.
What is Qwen3.8-Flash-Next's strongest current-benchmark score?
Terminal-Bench 2.1 at 86.1 (2026-08-28), with LiveBench at 76.2 next. The one headline number above them, GPQA Diamond 92.3, sits on a board this site grades saturated — kept on the record, not ranked.
Further reading
- GPQA Diamond leaderboard 2026 — Qwen3.8-Flash-Next is one of the 23 models it compares.
- GLM-5.3-Flash vs DeepSeek Flash vs Qwen3.8 — Qwen3.8-Flash-Next is one of the 3 models it compares.
- Models with 10M token context windows 2026 — Qwen3.8-Flash-Next is one of the 37 models it compares.