Meta (Meta Superintelligence Labs)
Muse Spark 1.2
All seven benchmark scores tracked here for Muse Spark 1.2 come from independent evaluators, not Meta's own runs — but on two of them, Meta's self-reported numbers run meaningfully higher: Terminal-Bench 2.1 (Meta 82.9% vs. Artificial Analysis's 80.15%, with vals.ai's own harness at just 69.66%) and DeepSWE v1.1 (Meta 59.3% vs. the independent leaderboard's 55%).
Muse Spark 1.2’s 7 benchmark scores on this page were verified against their sources on or after 2026-08-26.
- Released
- 2026-08-05
- License
- proprietary
- Context window
- 1M tokens
- Knowledge cutoff
- Not disclosed
The verified record
Against the 101 head-to-head comparisons Muse Spark 1.2 shares with other tracked models: 29 real gaps, 27 inside the noise band, and 45 we will not call.
A gap counts for Muse Spark 1.2 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Muse Spark 1.2 trails on 14 of them.
HLE · no toolsReasoning
±2 is noiseDeepSWELong-horizon coding
±9.5 is noiseLiveBenchComposite score across 7 domains
±2.7 is noiseToolathlon-VerifiedMulti-tool chores
±9.7 is noiseNo verdict for Muse Spark 1.2 anywhere on Terminal-Bench 2.1 (nothing independently confirmed on both sides); GPQA Diamond, SWE-bench Verified (saturated).
Muse Spark 1.2 API pricing
$1.25 in / $4.25 out per 1M tokens — official pricing
Muse Spark 1.2 benchmark scores
| Benchmark | Score |
|---|---|
| HLE(no tools)[1] Reasoning · ±2 is noise | |
| GPQA Diamondsaturated[2] Expert science Q&A — not ranked at any gap size | |
| Terminal-Bench 2.1[3] Terminal ops · ±10.6 is noise | |
| SWE-bench Verifiedsaturated[4] Bug fixing — not ranked at any gap size | |
| DeepSWE[5] Long-horizon coding · ±9.5 is noise | |
| Toolathlon-Verified[6] Multi-tool chores · ±9.7 is noise | |
| LiveBench[7] Composite score across 7 domains · ±2.7 is noise |
Who ran these numbers: 7 of 7 independent — artificialanalysis.ai (3), vals.ai (1), deepswe.datacurve.ai (1), toolathlon.xyz (1), livebench.ai (1).
- HLE: Reported at the model's "xhigh" reasoning-effort setting. Exact value (0.454587581093605) pulled from AA's embedded intelligence-index JSON payload on 2026-08-26; the rendered chart shows no per-bar text label. AA's own article rounds this to '44%'.
- GPQA Diamond: Reported at the model's "xhigh" reasoning-effort setting. Exact value (0.904040404040404) pulled from AA's embedded intelligence-index JSON on 2026-08-26. Not tested on vals.ai's GPQA Diamond leaderboard for this model (absent from its 20-benchmark coverage list and not found in vals.ai's top ~28 of 135 rows).
- Terminal-Bench 2.1: Reported at the model's "xhigh" reasoning-effort setting. Exact value (0.801498127340824) pulled from AA's embedded JSON. Meta's own launch chart (research.meta.ai) self-reports 82.9% on the same benchmark name; vals.ai's own harness (Mini-SWE-agent-style scaffold) scores it at only 69.66% — see background_facts for the three-way spread.
- SWE-bench Verified: Rank 13/86, ±1.52 CI, agent scaffold 'Mini-SWE-agent'. Page title/header confirm this is the SWE-bench Verified (500-instance) benchmark. No effort/tier suffix is shown on this leaderboard row.
- DeepSWE: Reported at the model's "xhigh" reasoning-effort setting. v1.1 leaderboard, rank 13, ±2% CI, avg cost $3.70, 99k output tokens, 101 steps. Meta self-reports 59.3% for the same benchmark on its own launch materials — a +4.3pp vendor-vs-independent gap.
- Toolathlon-Verified: Reported at the model's "xhigh" reasoning-effort setting. Pass@1 75.9 ±1.3 (rank 3 of 7 rows shown), Pass@3 87.0, Pass^3 63.0, 44.2 avg turns, 48.6 avg tool calls. Toolathlon-Verified release dated 2026-06-30; built by HKUST NLP (matches github.com/hkust-nlp/Toolathlon).
- LiveBench: Reported at the model's "xHigh Effort" setting. LiveBench-2026-06-25 release round (labelled 'latest' / live-updated). Category breakdown: Reasoning 90.0, Coding 77.5, Agentic Coding 57.6, Mathematics 91.2, Data Analysis 76.5, Language 78.6, Instruction Following 74.3, cost/successful task $0.375.
Notes on the record
Muse Spark 1.2 shipped 2026-08-05 as a coding-focused point release within Meta's Muse Spark line — the base model launched 2026-04-08, with Muse Spark 1.1 following on 2026-07-09 (as of 2026-08-26). It carries a 1,048,576-token (1M) context window under a proprietary license; knowledge cutoff is not disclosed in the data tracked here. Meta released it alongside Muse Code, its first terminal coding agent (beta); the model is available via the Meta Model API and through Muse Code at dev.meta.ai.
Pricing splits into two genuinely different tiers, not a simple discount. The Standard tier runs $1.25 in / $4.25 out per 1M tokens ($0.15 per 1M cache-hit tokens) and carries a no-training-use guarantee (developer.meta.com/ai/models/muse-spark/, as of 2026-08-26). The Contributor tier is $0.10 in / $0.20 out ($0.002 cache-hit) — roughly 12x/21x cheaper — but requires opting prompts and completions into Meta's training data. The cheap tier trades away the privacy guarantee; it isn't just a discount code.
Seven benchmark scores are tracked here, all sourced from independent evaluators rather than Meta's own runs, and mostly reported at the model's "xhigh" reasoning-effort setting (as of 2026-08-26): Humanity's Last Exam 45.46% (Artificial Analysis; AA's own article rounds this to "44%"), GPQA Diamond 90.4% (Artificial Analysis), Terminal-Bench 2.1 80.15% (Artificial Analysis), SWE-bench Verified 86.6% (vals.ai, rank 13/86, ±1.52 CI, Mini-SWE-agent scaffold — no effort tier is shown on this leaderboard row, so its comparability to the "xhigh" figures elsewhere is unconfirmed), DeepSWE v1.1 55% (deepswe.datacurve.ai, rank 13, ±2% CI), Toolathlon-Verified 75.9% Pass@1 (toolathlon.xyz, rank 3 of 7 rows shown; Pass@3 87.0%, Pass^3 63.0%), and LiveBench 78 (livebench.ai, 2026-06-25 release round; category breakdown: Reasoning 90.0, Mathematics 91.2, Coding 77.5, Data Analysis 76.5, Language 78.6, Instruction Following 74.3, Agentic Coding 57.6).
Two of those seven have a documented, precise vendor-vs-independent gap on record. On Terminal-Bench 2.1, Meta's own launch chart (research.meta.ai) self-reports 82.9%, versus Artificial Analysis's independent 80.15% (a +2.75pp vendor gap) and vals.ai's own harness at just 69.66% — a 13-point spread depending entirely on who ran the benchmark. On DeepSWE v1.1, Meta self-reports 59.3% versus the independent leaderboard's 55% — a +4.3pp vendor-favorable gap. Both are exact instances of the self-report/independent divergence this site exists to surface.
Coverage is not complete. GPQA Diamond has not been separately checked on vals.ai's own GPQA Diamond leaderboard — this model is absent from its 20-benchmark coverage list and wasn't found in the top ~28 of its 135 rows. Four other commonly tracked benchmarks were checked live on 2026-08-26 and have no row for this model at all: LiveCodeBench (vals.ai has zero Meta entries for this eval), ARC-AGI-2 (arcprize.org's newest Meta rows are still Llama 4 Scout/Maverick from April 2025), Agents' Last Exam (absent from all 39 rows of its leaderboard), and MathArena's HMMT Feb 2026 (absent from its ~31-model table). Separately, Artificial Analysis flags the model as comparatively verbose — 95M output tokens on its Intelligence Index versus a 72M median for similarly priced models, which matters directly for output-token cost under either pricing tier. Its overall AA Intelligence Index score is 57, ranked #17 of 187 tracked models (artificialanalysis.ai, as of 2026-08-26).
Licensing is a live, unresolved story. Muse Spark 1.2 shipped closed/proprietary on 2026-08-05. On 2026-08-10, Meta Chief AI Officer Alexandr Wang said open weights for Muse Spark 1.2 were "coming soon"; that same day Meta open-weighted a separate, smaller model, Muse Glimmer (30B, Apache 2.0) — not Muse Spark 1.2 itself. As of 2026-08-26, Muse Spark 1.2's weights remain unreleased: both Artificial Analysis ("isOpenWeights": false) and vals.ai ("WEIGHTS: PRIVATE") confirm proprietary status live. This site tracks that as an announced-but-unfulfilled openness promise, not settled fact.
Compare with
FAQ
Is Muse Spark 1.2 good at coding?
Coding scores vary sharply depending on who ran the test, not just on the model itself. On Terminal-Bench 2.1, Meta's own launch materials, Artificial Analysis's independent run, and vals.ai's own harness each report a different score for what is nominally the same benchmark — a 13-point spread from the lowest reading to the highest. Independent evaluators also score it well on SWE-bench Verified and DeepSWE, but this site treats any single coding number for Muse Spark 1.2 as harness-dependent rather than settled; see the tracked scores above for the full breakdown by source.
How does Muse Spark 1.2 perform on reasoning benchmarks like GPQA Diamond and Humanity's Last Exam?
Independent testing from Artificial Analysis puts Muse Spark 1.2 at 90.4% on GPQA Diamond, measured at the model's "xhigh" reasoning-effort setting. It also has a published Humanity's Last Exam score from the same source, though AA's own write-up rounds that figure down. GPQA Diamond doesn't have a second independent evaluator to cross-check against — it's absent from vals.ai's own GPQA Diamond leaderboard for this model — so treat that reading as single-source for now. (HLE's single-source status wasn't separately checked against other leaderboards.)
Is Muse Spark 1.2 open source?
No — Muse Spark 1.2 launched proprietary on 2026-08-05, and it still is as of 2026-08-26. Meta's Chief AI Officer Alexandr Wang said on 2026-08-10 that open weights were "coming soon," and Meta did open-weight a different, smaller model that same day (Muse Glimmer, Apache 2.0) — but not Muse Spark 1.2. As of 2026-08-26, both Artificial Analysis and vals.ai still list it as closed-weight, so treat the open-weighting promise as announced, not delivered.
How much does Muse Spark 1.2 cost to use via the API?
It depends which tier you pick, and the two aren't equivalent. The Standard tier costs $1.25 per 1M input tokens and $4.25 per 1M output tokens, and it carries a no-training-use guarantee. A separate Contributor tier is priced far lower, but it requires opting your prompts and completions into Meta's training data — that's a privacy trade, not just a cheaper plan. See the pricing details tracked on this page for the exact Contributor-tier rates.
What is Muse Code?
Muse Code is Meta's first terminal-based coding agent, released in beta alongside Muse Spark 1.2 on 2026-08-05 and available through dev.meta.ai. It's the delivery vehicle Meta is pairing with this model's coding-focused benchmark push (SWE-bench Verified, Terminal-Bench 2.1, DeepSWE, and Toolathlon-Verified are all tracked on this page), though this site does not yet have independent data on Muse Code's own performance separate from the underlying model.