Meta (Meta Superintelligence Labs)
Muse Spark 1.3
Four of the eleven benchmarks this site tracks already carry an independently-run Muse Spark 1.3 score one day after launch, and the clearest gain among them is LiveBench, +3.6 points at a matched reasoning-effort tier, clearing that benchmark's own noise band — Humanity's Last Exam's headline +3.6 points looks the same but mixes tiers against Muse Spark 1.2's own tracked score, shrinking to a marginal +2.04 on a fair same-tier comparison. DeepSWE has no independent score yet to check Meta's own newly self-reported 75.4% against.
Muse Spark 1.3 benchmarks and pricing, every number sourced: 4 tracked Muse Spark 1.3 benchmark scores (4 independently run, 0 still resting on a vendor’s own claim), priced at $1.25 per million input tokens and $4.25 per million output.
Muse Spark 1.3’s 4 benchmark scores on this page were verified against their source on 2026-09-03.
Version history: succeeded Muse Spark 1.2 (2026-08-05).
- Released
- 2026-09-02
- License
- proprietary
- Context window
- 1M tokens
- Knowledge cutoff
- Not disclosed
- Parameters
- Not disclosed
- Architecture
- Not disclosed
Muse Spark 1.3’s verified record
Against the 84 head-to-head comparisons Muse Spark 1.3 shares with other tracked models: 32 real gaps, 21 inside the noise band, and 31 we will not call.
A gap counts for Muse Spark 1.3 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Muse Spark 1.3 trails on 4 of them.
HLE · no tools Reasoning
±2 is noiseLiveBench Composite score across 7 domains
±2.7 is noiseNo verdict for Muse Spark 1.3 anywhere on Terminal-Bench 2.1 (every independently confirmed comparison inside the noise band); GPQA Diamond (saturated).
- No real gap yet in any of Muse Spark 1.3’s agentic comparisons — the independently confirmed ones all sit inside the noise band.
What changed from Muse Spark 1.2 to Muse Spark 1.3
The 4 benchmarks both models have been scored on, using the same variant each time. A raw Muse Spark 1.3 gain is not a real gain until it clears that benchmark’s own noise band, so each row below carries the verdict and not just the arithmetic.
| Benchmark | Muse Spark 1.2 | Muse Spark 1.3 | Change | Verdict |
|---|---|---|---|---|
| HLE | 45.46 | 49.1 | +3.6 | Real gap |
| Terminal-Bench 2.1 | 80.15 | 85.8 | +5.6 | Tie |
| GPQA Diamond | 90.4 | 93.8 | +3.4 | Tainted |
| LiveBench | 78 | 81.6 | +3.6 | Real gap |
Muse Spark 1.3 API pricing
$1.25 in / $4.25 out per 1M tokens — official pricing
What Muse Spark 1.3 costs per job
| Workload | Tokens in / out | Cost |
|---|---|---|
| One long chat turn | 100K / 10K | $0.168 |
| A codebase review | 1,000K / 100K | $1.68 |
| A day of agent work | 10,000K / 1,000K | $16.75 |
Computed from Muse Spark 1.3’s list rates above — cache discounts, its contributor tier (opts prompts/completions into meta's training data), and batch tiers are not applied.
Muse Spark 1.3 is one of 2 Meta (Meta Superintelligence Labs) models tracked on this site, at these official list prices.
| Model | In / 1M | Out / 1M |
|---|---|---|
| Muse Spark 1.3 | $1.25 | $4.25 |
| Muse Spark 1.2(superseded) | $1.25 | $4.25 |
Muse Spark 1.3 benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise | |
| Terminal-Bench 2.1[2] Terminal ops · ±10.6 is noise | |
| GPQA Diamondsaturated[3] Expert science Q&A — not ranked at any gap size | |
| LiveBench[4] Composite score across 7 domains · ±2.7 is noise |
Who ran these numbers: 4 of 4 independent — artificialanalysis.ai (3), livebench.ai (1).
- HLE: AA's own run at max effort, text-only subset (raw payload value 0.490732159406858). AA's xhigh-tier config for this model reads 47.5% instead — not tracked here.
- Terminal-Bench 2.1: AA's own run of 'Muse Spark 1.3 (max)' (0.857677902621723 in AA's payload — coincidentally identical to Gemini 3.7 Flash's own tracked figure; verified as a genuine tie, not a scraping duplicate, by confirming 229/267 against this benchmark's own 89-task x 3-trial structure). AA's xhigh-tier config reads 85.4% instead. Not yet on tbench.ai's official board, which lists only Muse Spark 1.1.
- GPQA Diamond: AA's own run at max effort (0.938383838383838 in AA's payload). Not on vals.ai's GPQA board, same as Muse Spark 1.2. AA's xhigh-tier config reads higher here, 94.14% — the one benchmark of the three where xhigh beats max for this model.
- LiveBench: Board row 'Muse Spark 1.3 xHigh Effort' on the 2026-06-25 LiveBench release — same release round and same xHigh tier as Muse Spark 1.1 and 1.2's own tracked rows.
Notes on the record
Muse Spark 1.3 launched September 2, 2026, per Meta's own developer page and AI Research blog post, independently corroborated by Bloomberg, Axios, and Neowin the same day. It follows the same roughly-monthly Muse Spark cadence tracked here before: base Muse Spark (2026-04-08), Muse Spark 1.1 (2026-07-09), Muse Spark 1.2 (2026-08-05), now 1.3. Meta's own materials describe it as trained on "months of broad adoption of Muse Code and Meta Model API" — the page does not state whether this is a new pretraining run or a post-training update to 1.2's checkpoint, unlike Google's explicit disclosure for Gemini 3.8 Flash.
Pricing is unchanged from Muse Spark 1.2: the Standard tier runs $1.25 in / $4.25 out per 1M tokens ($0.15 cache-hit) and carries a no-training-use guarantee; the Contributor tier is $0.10 in / $0.20 out ($0.002 cache-hit) but opts prompts and completions into Meta's training data — the same roughly 12x/21x discount-for-privacy trade this site has already flagged on 1.2's own page (confirmed side by side on developer.meta.com/ai/models/muse-spark/, checked 2026-09-03).
Artificial Analysis tracks two reasoning-effort configurations for 1.3: "max" (its new top tier, and the canonical/default page for this model) and "xhigh" (the same ceiling tier 1.2 was tracked at here). The two don't move in the same direction across benchmarks — max scores higher on Humanity's Last Exam (49.1% vs xhigh's 47.5%) and Terminal-Bench 2.1 (85.8% vs 85.4%), but xhigh scores higher on GPQA Diamond (94.1% vs max's 93.8%). This site tracks max as the default going forward, consistent with how Artificial Analysis presents the model's own canonical page — but that means the GPQA Diamond comparison against 1.2's own tracked 90.4% (itself an xhigh-tier figure) mixes tiers; on a matched xhigh-to-xhigh basis the improvement is larger (94.1% vs 90.4%). Of the four benchmarks 1.2 and 1.3 share here, only one — LiveBench — supports a clean, tier-matched, un-saturated real-vs-noise verdict: both models were scored at xHigh Effort, and the 3.6-point gain clears LiveBench's own noise band. Humanity's Last Exam's tracked 3.6-point gain (49.1% vs 45.46%) mixes tiers the same way GPQA Diamond's does — 1.3's tracked figure is AA's max-effort run, 1.2's is xhigh — and on a matched xhigh-to-xhigh basis (47.5% vs 45.46%) the gain shrinks to 2.04 points, barely above HLE's own 2.0-point noise floor rather than the comfortable margin the mixed-tier figure implies. Terminal-Bench 2.1's 5.6-point gain reads as a tie against that benchmark's own ±10.6 margin regardless of tier. GPQA Diamond's 3.4-point gain is moot either way — the benchmark is saturated and every comparison on it is tainted regardless of gap size.
Four of the eleven benchmarks this site tracks carry an independently-run Muse Spark 1.3 score, one day after launch: Humanity's Last Exam, Terminal-Bench 2.1, and GPQA Diamond (all Artificial Analysis), and LiveBench (81.6 overall on the 2026-06-25 release round, xHigh Effort tier — Reasoning 89.7, Coding 81.1, Agentic Coding 64.1, Mathematics 95.9, Data Analysis 79.6, Language 82.8, Instruction Following 78.0). Terminal-Bench 2.1's own canonical board, tbench.ai, still lists only Muse Spark 1.1, not 1.2 or 1.3.
Coverage gaps mirror 1.2's own, checked live on each as of 2026-09-03: LiveCodeBench (vals.ai has no Muse Spark row past 1.1), SWE-bench Verified (vals.ai's newest Meta row is still 1.2, at 86.60%), Toolathlon-Verified (toolathlon.xyz's newest Meta row is still 1.2, at 75.9%), DeepSWE (deepswe.datacurve.ai's newest Meta row is still 1.2's 55%, both on "Best" and "all effort levels" views), ARC-AGI-2, Agents' Last Exam, and HMMT Feb 2026 (all three have never carried any Muse Spark version, 1.3 included). Meta's own launch materials self-report 75.4% on DeepSWE v1.1 — a real gap from 1.2's own independently-verified 55% if the boards ever converge on 1.3, but that comparison cannot be made yet because no independent DeepSWE run of 1.3 exists.
Licensing is a harder story than 1.2's own "coming soon." Meta shipped 1.2 closed, with an August 10 promise from Chief AI Officer Alexandr Wang that open weights were imminent — still unfulfilled as of this writing. For 1.3, Meta has gone further: reporting describes the weights decision as explicitly undecided, citing the EU AI Act's Article 53 — the exemption for genuinely open-source models does not apply once a model is classified as carrying systemic risk, which frontier-scale models are liable to be. Wang separately claimed 1.3 is "competitive with Anthropic's Claude Fable 5.1, better than OpenAI's GPT-5.6 Sol at coding, and ahead of any current Chinese model" — a qualitative, self-reported cross-vendor comparison with no benchmark number attached in the reporting that carries it; this site records it as Meta's own claim, not a finding.
Compare with
FAQ
Has Muse Spark 1.3 been independently benchmarked?
Partially — four of the eleven benchmarks this site tracks carry an independently-run score as of 2026-09-03, one day after launch: GPQA Diamond (93.8%), Humanity's Last Exam (49.1%), and Terminal-Bench 2.1 (85.8%), all from Artificial Analysis at its "max" reasoning-effort tier, plus LiveBench (81.6 overall, xHigh Effort tier). LiveCodeBench, SWE-bench Verified, DeepSWE, ARC-AGI-2, Agents' Last Exam, and HMMT Feb 2026 have no Muse Spark 1.3 row on any board this site checked live — several of those have never scored any Muse Spark version at all, not just this one.
How does Muse Spark 1.3 pricing compare to Muse Spark 1.2?
It's identical: $1.25 per million input tokens and $4.25 per million output tokens on the Standard tier, confirmed side by side on Meta's own developer page. Meta also offers a cheaper Contributor tier at $0.10 in / $0.20 out — the same roughly 12x/21x discount as 1.2 — but that tier opts your prompts and completions into Meta's training data, unlike the privacy-preserving Standard tier.
Does Muse Spark 1.3 improve on Muse Spark 1.2's benchmark scores?
Only partly cleanly. LiveBench gains 3.6 points at a matched reasoning-effort tier (both models scored at xHigh Effort), clearing that benchmark's own noise band — a genuine improvement. Humanity's Last Exam's tracked score also gains 3.6 points, but that mixes tiers: 1.3's figure is Artificial Analysis's max-effort run, 1.2's is its xhigh-effort run; on a fair, matched xhigh-to-xhigh basis the gain shrinks to about 2.0 points, barely clearing that benchmark's own noise floor rather than the confident margin the mixed comparison suggests. Terminal-Bench 2.1 improves too, from 80.15% to 85.8%, but that 5.6-point gain sits inside this benchmark's own wide (±10.6) margin for error, so it counts as a tie. GPQA Diamond moves from 90.4% to 93.8%, but GPQA Diamond is saturated on this site (trust grade D) and its comparisons are tainted regardless of gap size. DeepSWE, where 1.2 showed the clearest self-report/independent gap on this site (Meta claimed 59.3%, the independent board measured 55%), has no independent score for 1.3 yet to check Meta's own newly claimed 75.4% against.
Is Muse Spark 1.3 open source?
No, and its path there is less certain than 1.2's. Muse Spark 1.2 shipped closed with an August 2026 promise from Meta's Chief AI Officer that open weights were coming soon — still unfulfilled. For 1.3, reporting describes Meta as having explicitly not decided whether to release weights, citing the EU AI Act's Article 53: the exemption it grants genuinely open-source models does not apply to a model classified as carrying systemic risk, a bar frontier-scale models are likely to meet regardless of licence.
Does Muse Spark 1.3 replace Muse Spark 1.2, or sit alongside it?
It replaces it as Meta's current release, though 1.2 (and 1.1) remain listed and available on Meta's own developer page rather than being withdrawn. This site marks Muse Spark 1.2 superseded accordingly — it stays on this site for comparison, not deleted, but is no longer the current recommendation.
Further reading
- Models with 10M token context windows 2026 — Muse Spark 1.3 is one of the 23 models it compares.