Alibaba

Qwen3.8-Flash-Next

All four benchmark numbers on this page — including the headline 73.5 Toolathlon-Verified and 58.7 DeepSWE 1.1 scores — are Alibaba's own self-reports; live checks of eight independent boards on 2026-08-26 found this two-day-old architecture preview listed on none of them.

Qwen3.8-Flash-Next’s 4 benchmark scores on this page were verified against their sources on or after 2026-08-26.

Released
2026-08-26
License
open-weights
Context window
262K tokens
Knowledge cutoff
Not disclosed

The verified record

Against the 55 head-to-head comparisons Qwen3.8-Flash-Next shares with other tracked models: 0 real gaps, 0 inside the noise band, and 55 we will not call.

A gap counts for Qwen3.8-Flash-Next only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Qwen3.8-Flash-Next trails on 0 of them.

No verdict for Qwen3.8-Flash-Next anywhere on Agents' Last Exam, DeepSWE, HLE · no tools, Toolathlon-Verified (nothing independently confirmed on both sides).

  • None of Qwen3.8-Flash-Next’s reasoning comparisons are independently confirmed on both sides yet.
  • None of Qwen3.8-Flash-Next’s coding comparisons are independently confirmed on both sides yet.
  • None of Qwen3.8-Flash-Next’s agentic comparisons are independently confirmed on both sides yet.

Qwen3.8-Flash-Next API pricing

in / out per 1M tokens official pricing

Qwen3.8-Flash-Next is one of 2 Alibaba models tracked on this site, at these official list prices.

Alibaba model pricing, official list rates
ModelIn / 1MOut / 1M
Qwen3.8-Flash-Next
Qwen3.8-Max$2.00$6.00

Qwen3.8-Flash-Next benchmark scores

Qwen3.8-Flash-Next benchmark scores, provenance, and source links
BenchmarkScore
HLE(no tools)[1]
Reasoning · ±2 is noise
DeepSWE[2]
Long-horizon coding · ±9.5 is noise
Toolathlon-Verified[3]
Multi-tool chores · ±9.7 is noise
Agents' Last Exam[4]
Professional work · ±3.2 is noise

Who ran these numbers: 0 of 4 independent; vendor self-reported (4).

  1. HLE: Reported for the model's "thinking mode" (default configuration). Vendor-reported in the HF model card benchmark table (footnote: 'judged by GPT-4o'), a different grading methodology than Artificial Analysis's own HLE runs. Confirmed live via artificialanalysis.ai's model-comparison search (searching 'flash-next' returns only DeepSeek V4 Flash Vision) that this model has no Artificial Analysis page yet, so no independent figure exists to prefer.
  2. DeepSWE: Card reports the higher of two harnesses (Claude Code vs. mini-SWE-agent; temp=1.0, top_p=0.95, 256K ctx) and states the model scores best on mini-SWE-agent -- the same harness the official board uses. Checked deepswe.datacurve.ai's v1.1 leaderboard live (last updated today, 2026-08-26): Qwen3.8-Flash-Next is not listed yet, only qwen3.8-max (57%±3%) appears for Alibaba. If later added, 58.7 would slot between claude-opus-4.8[max] (59%±2%) and qwen3.8-max[xhigh] (57%±3%) on that board.
  3. Toolathlon-Verified: Value is Pass@1. Vendor-reported Pass@1. Checked toolathlon.xyz's live Model Leaderboard (Toolathlon-Verified series): Qwen3.8-Flash-Next is not listed as of today; the independent top score there is Kimi K3 (max) at 76.5±1.9, and the only Qwen entry present is Qwen3.5 397B-A17B (40.7±2.0, dated 2026-07-30).
  4. Agents' Last Exam: Value is Pass@1 (a separate composite "Score" of 51.2 is also reported by the vendor but not used here). Vendor card reports both Pass@1 (24.3) and a separate composite 'Score' (51.2); Pass@1 used here since snorkel.ai's live leaderboard ranks primarily by Pass Rate. Checked snorkel.ai/leaderboard/agents-last-exam/ live: Qwen3.8-Flash-Next does not appear in the model filter dropdown as of today (only Qwen3 6-Plus, Qwen3 7-Max, Qwen3 8-27B, Qwen3 8-Max are listed for Alibaba).

Notes on the record

Qwen3.8-Flash-Next shipped as an open-weight-only architecture preview, not a production release (Hugging Face repo metadata: created 2026-08-24, last modified 2026-08-26). Despite the "Qwen3.8" version label, Hugging Face's own architecture metadata lists it as "Qwen4ExpForConditionalGeneration" (model_type "qwen4_exp") — the card frames it as an experimental preview of the architecture that will underpin Qwen4, not a Qwen3.8-generation model in the strict sense. No hosted API price exists for this exact checkpoint as of 2026-08-26 — no input or output price is listed anywhere, and a live check found no DashScope or Model Studio listing. The $0.16/1M input and $0.47/1M output figures circulating online belong to a different, not-yet-shipped sibling product: the bare "Qwen3.8-Flash" (without "-Next"), which the HF model card itself describes as "the official version based on Qwen3.8-Flash-Next with more production features, e.g., 1M context length by default, official built-in tools," to be served via Qwen Cloud (confirmed by a Qwen/Alibaba_Qwen X post, which is also the source of that pricing). This page's 262,144-token context window is the -Next preview's NATIVE length; the card separately says the preview is "extensible up to 1,000,000 tokens" via RoPE-scaling techniques such as YaRN, which is a different claim from the sibling's "1M context length by default" — extensible-with-work versus shipped-as-standard. Neither figure should be read as the other's. Licensing is "qwen-community-1.0" (Qwen Community License 1.0) per Hugging Face's structured API — a custom license granting broad commercial-use rights, distinct from the OSI-approved Apache-2.0 template; the same API confirms the checkpoint is gated:false, disabled:false, across 131 safetensors shards. Those shards also resolve an apparent size contradiction: Alibaba's headline "125B total, 6B active" describes only the Mixture-of-Experts core, while the full on-disk checkpoint totals roughly 180B parameters (179,999,981,459 across all dtypes — 179,999,981,424 of them BF16, per Hugging Face's safetensors metadata) once a 51B-parameter N-gram embedding layer and a 4B-parameter multi-token-prediction head are added in — both figures are accurate, they just answer different questions.

All four benchmark scores on this page are self-reported by Alibaba (source: the Hugging Face model card's benchmark table), and none has independent confirmation as of 2026-08-26. Live checks that day found the model listed on none of eight independent boards: artificialanalysis.ai (whose own search for "flash-next" surfaces only DeepSeek V4 Flash Vision), vals.ai, arcprize.org, livebench.ai, matharena.ai, deepswe.datacurve.ai (updated that same day), toolathlon.xyz, and snorkel.ai's Agents' Last Exam model filter. LiveBench in particular cannot add it soon on structural grounds: its snapshot cadence is six months and its current live release is dated 2026-06-25, two months before this model existed. The vendor numbers themselves also carry methodology flags worth noting: the 35.9 HLE score (thinking mode, the model's default configuration) is, per the card's own footnote, "judged by GPT-4o" — a different judge than Artificial Analysis uses in its independent HLE runs; the 58.7 DeepSWE 1.1 figure is the higher of two harnesses (Claude Code vs. mini-SWE-agent), and the card notes the model does best specifically on mini-SWE-agent — the same harness the official DeepSWE board itself uses, which is why a slot between claude-opus-4.8[max] (59%±2%) and qwen3.8-max[xhigh] (57%±3%) is a reasonable placement if the board, last updated the same day but still without this model, ever adds it; and the 24.3 Agents' Last Exam figure is Pass@1, chosen here over the vendor's separate 51.2 composite "Score" because Snorkel AI's own board ranks primarily by Pass Rate. All four numbers should be read as vendor-claimed pending independent re-confirmation on this exact checkpoint, not as settled scores.

Compare with

FAQ

What is Qwen3.8-Flash-Next actually good at, based on the available numbers?

Alibaba's own tables show strength in agentic coding and tool use: 58.7 on DeepSWE 1.1 (the better of two test harnesses — mini-SWE-agent, the same harness the official DeepSWE board itself uses) and 73.5 Pass@1 on Toolathlon Verified. But every one of these figures is vendor-reported — no independent leaderboard has run this model yet, so its real-world standing against verified scores like Kimi K3 [max]'s 76.5±1.9 on Toolathlon is still unknown.

How much does Qwen3.8-Flash-Next cost to use via API?

The price you may have seen quoted — $0.16 per 1M input tokens, $0.47 per 1M output tokens — is not this model's price; that figure was posted by Alibaba for a separate, not-yet-released sibling product, the bare "Qwen3.8-Flash" (without "-Next"), which will reportedly add 1M-token context and built-in tools once it ships on Qwen Cloud. As of 2026-08-26, Qwen3.8-Flash-Next itself has no DashScope or Model Studio listing at all.

Why doesn't this model show up on any independent benchmark site?

The model is only about two days old as of the data collection date, and eight boards checked on 2026-08-26 — Artificial Analysis, Vals AI, ARC Prize, LiveBench, MathArena, DeepSWE's own leaderboard, Toolathlon, and Snorkel AI's Agents' Last Exam board — had none added it yet. One of those, LiveBench, runs on a much slower cycle than the rest: it only takes a new snapshot roughly twice a year, and its most recent one is dated before this model even existed, so it won't catch up quickly regardless of attention.

Is Qwen3.8-Flash-Next open source?

It is open-weights, released under a custom license called qwen-community-1.0 (Qwen Community License 1.0) — a permissive license granting broad commercial-use rights, distinct from the OSI-approved Apache-2.0 template. Hugging Face's structured metadata also lists the checkpoint as ungated and not disabled.

Why do different sources list this model's size as 6B, 125B, or 180B parameters?

None of the three numbers is wrong — each measures a different slice of the same checkpoint. The 6B and 125B figures both describe the Mixture-of-Experts core: 6B is what's active per token, 125B is that core's full size. The 180B figure is the entire on-disk checkpoint, which stacks two more components on top of that core — a large embedding layer used for cheap parameter scaling and a smaller head that predicts multiple tokens at once. Hugging Face's safetensors metadata puts the grand total at 179,999,981,459 parameters, all but 35 of them BF16.

Get the next verdict by email

One email per verdict — which launch-chart claims held up. No spam, unsubscribe anytime.