No one sells a real 10M-token window — and five of nineteen models charge more before you reach the max

August 28, 2026

"Models with 10 million token context window" is tied for the highest-volume single spec query in this site's keyword research. Here is the answer nobody selling one will give you: of the 19 models tracked here, none ships a 10M-token window — native, extended, or otherwise. The biggest real number on sale is 1,050,000 tokens, and across the 19 tracked models the spec sheet splits three ways: 13 models charge one flat rate all the way to their maximum, 5 quietly move you to a higher price once your prompt crosses a threshold the marketing page doesn't lead with, and 1 ships a smaller native window that only reaches 1M if you re-configure it yourself. Every number below was checked against the vendor's own documentation on 2026-08-28, and each table names the vendor pages it came from.

Where the 10M number actually comes from

A context-window figure can mean two different things, and the gap between them is where the biggest claims live. A native window is what the model was trained to handle and what the default configuration serves. An extended window is what you can reach by stretching the model's position encoding — techniques like RoPE scaling and YaRN — which the vendor may support, recommend, or merely permit, and which the model was not trained at.

The cleanest worked example on this site is Qwen3.8-Flash-Next. Its Hugging Face card states a native length of 262,144 tokens, "extensible up to 1,000,000 tokens" via YaRN — the vendor's own words, and an honest framing. The same card describes a separate sibling product — the bare Qwen3.8-Flash — with "1M context length by default," and that sibling went on sale through Alibaba's cloud API on 2026-08-26 (OpenRouter lists the live maker-operated endpoint at a 1,000,000-token window; we found no first-party Alibaba announcement page). Those are three different claims — 262K trained on this checkpoint, 1M reachable on it with work, and 1M standard on a different product — and a spec-table that prints "1M" for this model collapses all three into one number.

That is the mechanism behind every headline-grabbing 10M figure we could find: a length-extension claim, not a trained length. No model tracked here makes even the extension claim at 10M, so this table stops at what is actually on sale.

The five models where 1M costs more than the price page led with

Five tracked models publish a per-token rate, and then a second, higher rate that kicks in at or above a prompt-length threshold. The step isn't a secret — the thresholds and both rates below are from each vendor's own pricing page — but every one of them quotes the lower rate first. If your workload lives in long context, the second column is your real price.

ModelWindowThresholdBelow → above (per 1M tokens)Source, checked 2026-08-28
GPT-5.6 Sol1,050,000>272K input$4.00/$20.00 → $8.00/$30.00 (2× in, 1.5× out)developers.openai.com
GPT-5.6 Luna1,050,000>272K input$0.20/$1.20 → $0.40/$1.80 (same family rule)developers.openai.com
Gemini 3.1 Pro Preview1,048,576>200K prompt$2.00/$12.00 → $4.00/$18.00ai.google.dev pricing
Grok 4.6500,000≥200K prompt$2.00/$6.00 → $4.00/$12.00 (2× everything)docs.x.ai
MiniMax M31,000,000>512K input$0.30/$1.20 → $0.60/$2.40 (all three rates step up)minimax.io model page + pricing guide

Three details worth reading twice. Grok 4.6's boundary is inclusive — xAI's own row label reads "≥ 200k", so a prompt of exactly 200,000 tokens already bills at the doubled rate. MiniMax's table steps up all three of its rates — input, output, and cache reads — for requests above 512K input; unlike xAI and OpenAI, its page never states whether tokens below the line keep the lower rate. And OpenAI's rule applies "for the full request" too: crossing 272K doesn't split the bill, it doubles the input side of all of it.

One more pricing footnote: two of the five base rates above are themselves promotional — OpenAI labels GPT-5.6 Sol's $4.00/$20.00 a promo guaranteed at least through 2026-11-21, and MiniMax badges its $0.30/$1.20 as a "Permanent 50% off" rate against list prices of $0.60/$2.40 and $1.20/$4.80 per tier.

None of this makes the windows fake. It makes the phrase "1M context window at $2 per million tokens" incomplete — the window is real, the price applies to less of it than the sentence implies.

The thirteen that charge one rate to the ceiling

The other 13 models with published pricing charge the same per-token rate at token 1,000,000 as at token 1,000 — confirmed against each vendor's own pricing page, not assumed from the absence of a disclosed tier.

ModelWindow (vendor-stated)Flat-rate note, checked 2026-08-28
Claude Fable 51,000,000Anthropic bills a 900K-token request at the same per-token rate as a 9K one (its own docs' example, Claude 4.6+); speed and batch variants exist, none keyed to length
Claude Opus 51,000,000Same rule
Claude Opus 4.81,000,000Same rule
Claude Sonnet 51,000,000Same rule
DeepSeek V4 Pro (0813)1,000,000Rates vary by time of day, never by prompt length; output capped at 384K
DeepSeek V4 Flash (0731)1,000,000Same peak/off-peak structure, no length tier
Kimi K31,048,576Single rate on Moonshot's pricing page
GLM-5.31,000,000 ("1M")Single rate on docs.z.ai
GLM-5.3-Flash1,048,576 ("1M")Single rate; launch promo aside, no length tier
GLM-5.21,000,000 ("1M")Single rate
Qwen3.8-Max1,000,000One rate per region — Singapore bills above the other five — but no length tier
Muse Spark 1.21,048,576 ("1M")No length tier on Meta's pricing table
Gemini 3.7 Flashup to 1,048,576The explicit exception in Google's own lineup: no length-based price split, unlike its Pro sibling

Two attribution notes, because precision is the point of this page: several vendors' pages say only "1M" — where they do, the table marks it ("1M") and the integer shown follows that model's own config or API reference, which is why the same label renders as 1,000,000 on some rows and 1,048,576 on others. And Gemini 3.7 Flash's card lists 1M with no length-based price split — Google's own stated structure for that model, sitting one product away from the Pro model that does carry a 200K split.

Qwen3.8-Flash-Next sits outside both tables: it has no first-party API to price at all — 262,144 tokens native as open weights, 1M reachable via YaRN if you serve it yourself.

How to read a context-window claim now

Three questions turn any context-window spec into something you can actually budget against. Is it native or extended? — if the card says "extensible," the headline number was not trained, and the two claims are recorded separately here. Is the rate flat across the window? — five of nineteen models here say no once you find the right table row, and one of them starts charging double at exactly the threshold, not past it. Who says so, and when? — every claim above carries the vendor page it came from and the date we read it, which is also why this piece publishes no blended "context quality score": a composite number would hide exactly the distinctions this table exists to show. The per-model records — pricing, scores, and what changed since — update on each model's own page before they appear anywhere else; start from the full catalog.


Sources: every context-window figure, threshold, and rate above is from the named vendor's own documentation — Anthropic's model comparison table, developers.openai.com model pages, ai.google.dev's pricing page and the Gemini 3.7 Flash model card, docs.x.ai's model table, MiniMax's own M3 model page and platform.minimax.io's pricing guide, the Gemini API models table (source of the 3.1 Pro window figure), api-docs.deepseek.com's pricing page, Moonshot's chat-k3 docs, docs.z.ai's pricing pages, Alibaba Cloud Model Studio's qwen3-8-max page, developer.meta.com's Muse Spark page, and the Qwen/Qwen3.8-Flash-Next Hugging Face card — each read on 2026-08-28. Per-model provenance for benchmark scores lives on the linked model pages.