StepFun
Step 5 Preview
StepFun pitches Step 5 Preview as frontier-level for software engineering at $1.00/$2.70 per million tokens, but only one of its six tracked benchmark scores — Humanity's Last Exam without tools, 46.48, run by Artificial Analysis — is independent. The coding and agent numbers behind the pitch are StepFun's own, and none of the coding or agent boards this site tracks lists the model yet.
Step 5 Preview benchmarks and pricing, every number sourced: 6 tracked Step 5 Preview benchmark scores (1 independently run, 5 still resting on a vendor’s own claim), priced at $1.00 per million input tokens and $2.70 per million output.
Step 5 Preview is Mixture-of-Experts with 600B total parameters (27B activated per token) and a 1M-token context window.
Step 5 Preview’s 6 benchmark scores on this page were verified against their source on 2026-09-22.
- Released
- 2026-09-20
- License
- proprietary
- Context window
- 1M tokens
- Knowledge cutoff
- Not disclosed
- Parameters
- 600B (27B active)
- Architecture
- Mixture-of-Experts
Step 5 Preview’s verified record
Step 5 Preview’s most-compared rival is DeepSeek V4 Pro (0813): 1 lead and 5 not callable across their 6 shared comparisons. Step 5 Preview is priced at $1.00/$2.70 per 1M tokens (in/out) vs DeepSeek V4 Pro (0813)’s $1.32/$3.96.
Against the 124 head-to-head comparisons Step 5 Preview shares with other tracked models: 21 real gaps, 6 inside the noise band, and 97 we will not call.
A gap counts for Step 5 Preview only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Step 5 Preview trails on 9 of them.
HLE · no tools Reasoning
±2 is noiseNo verdict for Step 5 Preview anywhere on DeepSWE, HLE · with tools, Terminal-Bench 2.1, Toolathlon-Verified (nothing independently confirmed on both sides); GPQA Diamond (saturated).
- None of Step 5 Preview’s coding comparisons are independently confirmed on both sides yet.
- None of Step 5 Preview’s agentic comparisons are independently confirmed on both sides yet.
Step 5 Preview API pricing
$1.00 in / $2.70 out per 1M tokens — official pricing
What Step 5 Preview costs per job
| Workload | Tokens in / out | Cost |
|---|---|---|
| One long chat turn | 100K / 10K | $0.127 |
| A codebase review | 1,000K / 100K | $1.27 |
| A day of agent work | 10,000K / 1,000K | $12.70 |
Computed from Step 5 Preview’s list rates above — cache discounts and batch tiers are not applied.
Step 5 Preview benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise | |
| HLE(with tools)[2] Reasoning · ±2 is noise | |
| GPQA Diamondsaturated[3] Expert science Q&A — not ranked at any gap size | |
| Terminal-Bench 2.1[4] Terminal ops · ±10.6 is noise | |
| DeepSWE[5] Long-horizon coding · ±9.5 is noise | |
| Toolathlon-Verified[6] Multi-tool chores · ±9.7 is noise |
Who ran these numbers: 1 of 6 independent — artificialanalysis.ai (1); vendor self-reported (5).
- HLE: Artificial Analysis's own run (text-only 2,158-question subset, pass@1; the board rounds this to 46.5). AA lists the model as "Step 5 Preview" with no effort level, so which reasoning tier produced the score is not stated. StepFun's launch table reports the same 46.5, at its High setting.
- HLE: StepFun's own launch-table figure (High effort). The vendor's footnote says Step 5 Preview and GLM-5.3 were evaluated on the text-only subset while every other model in its table used the full dataset, and that the settings are not directly comparable. This site has found no independent with-tools run.
- GPQA Diamond: StepFun's own launch-table figure (High effort; harness not stated). GPQA Diamond is graded saturated on this site, so the number carries little weight either way; vals.ai no longer runs new models on its GPQA page, and Artificial Analysis had no Step 5 Preview entry there (checked 2026-09-21).
- Terminal-Bench 2.1: StepFun's own launch-table figure, labeled Terminal-Bench v2.1 (High effort); the agent harness is not stated. tbench.ai's 2.1 board (last updated 2026-09-03) does not list the model. Artificial Analysis now scores Terminal-Bench 4.0 instead, where it reports 33.3 (the same figure StepFun shows on its own Terminal-Bench v4 row) — a different benchmark this site does not track.
- DeepSWE: StepFun's own launch-table figure (High effort), run on the SWE-agent harness at temperature 1.0 and top_p 0.95. The independent deepswe.datacurve.ai board runs every model on mini-swe-agent, so this is not a like-for-like number, and that board (v1.1, updated 2026-09-03) does not list the model.
- Toolathlon-Verified: StepFun's own launch-table figure (High effort; harness not stated). toolathlon.xyz's board (latest entry 2026-08-30) does not list the model.
Notes on the record
StepFun announced Step 5 Preview on 2026-09-20 as its flagship for agentic work; it is hosted-only (products and API, no downloadable weights). Artificial Analysis had listed it, with a 2026-09-18 release label, at least 21 hours earlier. Pricing on StepFun's international USD page (checked 2026-09-21): $1.00 per million input tokens on a cache miss, $0.05 on a cache hit and $2.70 for output, reasoning tokens included. It is one flat rate, with no length, time-of-day or promotional condition anywhere in StepFun's docs. Context window is 1M tokens; maximum output is 64K tokens — the docs listed 1M at launch and were edited within two days.
Per StepFun, it is a sparse mixture-of-experts with 600B total and 27B active parameters, taking text, images and video and returning text. Its docs offer three reasoning-effort levels (low, medium, high); every StepFun figure on this page is at High, the top level. StepFun says weights will be released on October 15 — a plan, with no year, license or repository stated — so this site lists it as proprietary until they exist. Unofficial copies labeled Step-5-Preview-BF16 have appeared on Hugging Face since 2026-09-20; StepFun's own repository is not publicly visible (checked 2026-09-22), and StepFun's site, docs and X account say nothing about the copies.
One of the six benchmark scores tracked here is independent: Humanity's Last Exam without tools, 46.48, from Artificial Analysis's text-only run. AA states no effort level for that run, so the record's HLE verdicts set a run of unstated tier against rivals' top-tier runs. The other five come from StepFun's launch table. Its with-tools HLE figure uses the text-only subset while most rivals' figures use the full dataset, and its DeepSWE run used the SWE-agent harness, not the mini-swe-agent harness the independent board applies to every model. StepFun's Agents' Last Exam figure is the harder ALE-CLI split, not the split this site tracks, so it is not recorded.
Compare with
FAQ
Has Step 5 Preview been independently benchmarked?
Partly. Artificial Analysis has run it on Humanity's Last Exam (46.48, text-only, no effort level shown), the only one of this site's eleven tracked benchmarks with an independent score for this model. StepFun's own launch table supplies GPQA Diamond (93.5), HLE with tools (59.4), Terminal-Bench 2.1 (85.0), DeepSWE (67.7) and Toolathlon-Verified (74.1), all self-reported. As of 2026-09-21 the boards for DeepSWE, Terminal-Bench 2.1 (tbench.ai), Toolathlon, Agents' Last Exam, ARC Prize, LiveBench and MathArena did not list the model; Artificial Analysis also scores it on Terminal-Bench 4.0 (33.3), which this site does not track.
Is Step 5 Preview open source?
Not yet. StepFun says the weights will be released on October 15 but has given no year, license or repository, and its official Hugging Face repository is not publicly visible. Third-party copies labeled Step-5-Preview-BF16 have been posted since 2026-09-20; they are not a StepFun release, StepFun's site, docs and X account say nothing about them, and StepFun has named no license, so treat any license label on a copy as unverified. Until StepFun publishes weights, this site lists the model as proprietary and hosted-only.
What does Step 5 Preview cost?
StepFun's international USD price list (checked 2026-09-21) charges $1.00 per million input tokens on a cache miss, $0.05 on a cache hit and $2.70 per million output tokens, with reasoning tokens billed as output. It is a single flat rate: no length threshold, time-of-day pricing or promotional price appears in StepFun's docs. StepFun's China-region site lists the same model in yuan (¥7, ¥0.35 and ¥20 per million tokens).
How does Step 5 Preview compare with the frontier on coding?
There is no independent evidence on the coding benchmarks this site tracks yet. StepFun's own launch table puts its DeepSWE score (67.7) within that benchmark's 9.5-point noise band of Kimi K3 (67.5), GLM-5.3 (66.9), GPT-6 Astra (74.1) and Claude Opus 5 (74.0) — none of those distances would count as a gap here even if independently run — and the table is StepFun's, and its DeepSWE run used the SWE-agent harness, not the mini-swe-agent the independent DeepSWE board applies to every model. That board does not list the model, so this site calls none of those gaps. Artificial Analysis does report 33.3 on Terminal-Bench 4.0, a newer benchmark this site does not track.
What can Step 5 Preview take as input, and how much can it write?
Text, images and video in; text out. The context window is 1M tokens and the maximum output is 64K tokens — StepFun's docs listed 1M for maximum output at launch and changed it to 64K within two days. Reasoning effort can be set to low, medium or high; StepFun's benchmark figures use high.
Did Step 5 Preview leak before StepFun announced it?
Not the weights, on the evidence this site could check. Artificial Analysis had a page for the model, with results, a price and a 2026-09-18 release label, in an archive capture from 05:44 UTC on 2026-09-19 — at least 21 hours before StepFun's announcement post. AA's page does not say when it was given access, and this site does not read that gap as a leak. The first of several unofficial copies labeled as its weights appeared on Hugging Face about two hours after the post, at 05:07 UTC on 2026-09-20. StepFun's site, docs and X account say nothing about either, and this site found no anonymous OpenRouter, OpenCode or Arena listing tied to the model.
Further reading
- Models with 10M token context windows 2026 — Step 5 Preview is one of the 31 models it compares.