Meta (Meta Superintelligence Labs)Previous version · Muse Spark 1.3

Muse Spark 1.2

All eight benchmark scores tracked here for Muse Spark 1.2 come from independent evaluators, not Meta's own runs — but independent sources don't agree with each other either: Terminal-Bench 2.1 spans a 13-point range across Meta's own chart (82.9%), Artificial Analysis (80.15%), and vals.ai's harness (69.66%), and on DeepSWE v1.1 Meta's self-reported 59.3% runs a real 4.3-point above the independent leaderboard's 55%, outside that benchmark's own ±2% CI.

Muse Spark 1.2 benchmarks and pricing, every number sourced: 8 tracked Muse Spark 1.2 benchmark scores (8 independently run, 0 still resting on a vendor’s own claim), priced at $1.25 per million input tokens and $4.25 per million output.

Released
2026-08-05
License
proprietary
Context window
1M tokens
Knowledge cutoff
Not disclosed
Verified
sources checked 2026-08-26–2026-10-08
Parameters
Not disclosed
Architecture
Not disclosed

Muse Spark 1.2’s verified record

Muse Spark 1.2’s most-compared rival is DeepSeek V4 Pro (0813): 1 lead, 2 ties, and 4 not callable across their 7 shared comparisons. Muse Spark 1.2 is priced at $1.25/$4.25 per 1M tokens (in/out) vs DeepSeek V4 Pro (0813)’s $1.32/$3.96.

Against the 183 head-to-head comparisons Muse Spark 1.2 shares with other tracked models: 24 real gaps, 24 inside the noise band, and 135 we will not call.

A gap counts for Muse Spark 1.2 only where both sides were run independently and the benchmark still separates models. Losses are listed alongside wins on purpose — Muse Spark 1.2 trails on 9 of them.

HLE · no tools Reasoning

±2 is noise
Behind
MiMo-V2.6-Pro −3.9 · 1 previous: Claude Opus 4.8 −3.2
Ahead
Qwen3.8-Max +2.4 · Grok 4.7 +2.4 · DeepSeek V4 Pro (0813) +4.5 · GLM-5.3-Flash +5.6 · MiniMax M3 +6.5 · Qwen3.8-Flash-Next +7.5 · 1 previous: GLM-5.2 +4.4
Tie
3 models within ±2
Unverified
1 model — vendor-reported on one side
Setup-dependent
21 models — scored on a different harness or effort tier

LiveBench Composite score across 7 domains

±2.7 is noise
Behind
Ahead
Claude Haiku 5.5 +5.9 · GLM-5.3-Flash +6.4 · MiniMax M3 +10.7 · 2 previous: DeepSeek V4 Flash (0731) +3.8 · GLM-5.2 +4.8
Tie
9 models within ±2.7
Setup-dependent
14 models — scored on a different harness or effort tier

DeepSWE Long-horizon coding

±9.5 is noise
Behind
Kimi K3 −14.0 · GPT-6 Astra −19.0 · 3 previous: Grok 4.6 −12.0 · Claude Fable 5 −15.0 · Claude Opus 5 −19.0
Ahead
1 previous: GLM-5.2 +11.0
Tie
4 models within ±9.5
Unverified
11 models — vendor-reported on one side
Setup-dependent
8 models — scored on a different harness or effort tier

Toolathlon-Verified Multi-tool chores

±9.7 is noise
Ahead
Gemini 3.1 Pro Preview +14.8 · 1 previous: GLM-5.2 +16.0
Tie
6 models within ±9.7
Unverified
8 models — vendor-reported on one side

Terminal-Bench 4.0 · xhigh Terminal ops

±12.4 is noise
Behind
Grok 4.7 −22.7

No verdict for Muse Spark 1.2 anywhere on Terminal-Bench 2.1 ( every independently confirmed comparison inside the noise band); GPQA Diamond, SWE-bench Verified ( saturated).

Muse Spark 1.2 API pricing

$1.25 in / $4.25 out per 1M tokens — official pricing source

What Muse Spark 1.2 costs per job

Muse Spark 1.2 cost for three reference workloads, computed from its list rates
WorkloadTokens in / outCost
One long chat turn100K / 10K$0.168
A codebase review1,000K / 100K$1.68
A day of agent work10,000K / 1,000K$16.75

Computed from Muse Spark 1.2’s list rates above — cache discounts, its contributor tier (opts prompts/completions into meta's training data), and batch tiers are not applied.

Muse Spark 1.2 is one of 2 Meta (Meta Superintelligence Labs) models tracked on this site, at these official list prices.

Meta (Meta Superintelligence Labs) model pricing, official list rates
ModelIn / 1MOut / 1M
Muse Spark 1.3$1.25$4.25
Muse Spark 1.2(previous version)$1.25$4.25

Muse Spark 1.2 benchmark scores

Muse Spark 1.2 benchmark scores, provenance, and source links
BenchmarkScore
HLE(no tools)[1]
Reasoning · ±2 is noise
GPQA Diamondsaturated[2]
Expert science Q&A — not ranked at any gap size
Terminal-Bench 2.1[3]
Terminal ops · ±10.6 is noise
SWE-bench Verifiedsaturated[4]
Bug fixing — not ranked at any gap size
DeepSWE[5]
Long-horizon coding · ±9.5 is noise
Toolathlon-Verified[6]
Multi-tool chores · ±9.7 is noise
LiveBench[7]
Composite score across 7 domains · ±2.7 is noise
Terminal-Bench 4.0(xhigh)[8]
Terminal ops · ±12.4 is noise

Who ran these numbers: 8 of 8 independent — artificialanalysis.ai (3), vals.ai (2), deepswe.datacurve.ai (1), toolathlon.xyz (1), livebench.ai (1).

  1. HLE: Reported at the model's "xhigh" reasoning-effort setting. Exact value pulled from AA's embedded intelligence-index JSON payload on 2026-08-26; the rendered chart shows no per-bar text label. AA's own article rounds this to '44%'.
  2. GPQA Diamond: Reported at the model's "xhigh" reasoning-effort setting. Exact value pulled from AA's embedded intelligence-index JSON on 2026-08-26. Not tested on vals.ai's GPQA Diamond leaderboard for this model (absent from its 20-benchmark coverage list and not found in vals.ai's top ~28 of 135 rows).
  3. Terminal-Bench 2.1: Reported at the model's "xhigh" reasoning-effort setting. Exact value pulled from AA's embedded JSON. Meta's own launch chart (research.meta.ai) self-reports 82.9% on the same benchmark name; vals.ai's own Terminus 2 run scores it at 69.66%.
  4. SWE-bench Verified: Rank 13/86, ±1.52 CI, agent scaffold 'Mini-SWE-agent'. Page title/header confirm this is the SWE-bench Verified (500-instance) benchmark. No effort/tier suffix is shown on this leaderboard row.
  5. DeepSWE: Reported at the model's "xhigh" reasoning-effort setting. v1.1 leaderboard, rank 13, ±2% CI, avg cost $3.70, 99k output tokens, 101 steps. Meta self-reports 59.3% for the same benchmark on its own launch materials — a +4.3pp vendor-vs-independent gap.
  6. Toolathlon-Verified: Reported at the model's "xhigh" reasoning-effort setting. Pass@1 75.9 ±1.3 (rank 3 of 7 rows shown), Pass@3 87.0, Pass^3 63.0, 44.2 avg turns, 48.6 avg tool calls. Toolathlon-Verified release dated 2026-06-30; built by HKUST NLP (matches github.com/hkust-nlp/Toolathlon).
  7. LiveBench: Reported at the model's "xHigh Effort" setting. LiveBench-2026-06-25 release round (labelled 'latest' / live-updated). Category breakdown: Reasoning 90.0, Coding 77.5, Agentic Coding 57.6, Mathematics 91.2, Data Analysis 76.5, Language 78.6, Instruction Following 74.3, cost/successful task $0.375.
  8. Terminal-Bench 4.0: vals.ai board row — mini-swe-agent harness (single bash tool), pass@1 averaged over 3 full passes, raw value 6.061 at effort xhigh. No official or Artificial Analysis Terminal-Bench 4.0 row exists for Muse Spark 1.2.

Notes on the record

Muse Spark 1.2 shipped 2026-08-05 as a coding-focused point release within Meta's Muse Spark line — the base model launched 2026-04-08, with Muse Spark 1.1 following on 2026-07-09 (as of 2026-08-26). It carries a 1,048,576-token (1M) context window under a proprietary license; knowledge cutoff is not disclosed in the data tracked here. Meta released it alongside Muse Code, a terminal coding agent (beta); the model is available via the Meta Model API and through Muse Code at dev.meta.ai.

Pricing splits into two genuinely different tiers, not a simple discount. The Standard tier runs $1.25 in / $4.25 out per 1M tokens ($0.15 per 1M cache-hit tokens) and carries a no-training-use guarantee (developer.meta.com/ai/models/muse-spark/, as of 2026-08-26). The Contributor tier is $0.10 in / $0.20 out ($0.002 cache-hit) — roughly 12x/21x cheaper — but requires opting prompts and completions into Meta's training data. The cheap tier trades away the privacy guarantee; it isn't just a discount code.

Eight benchmark scores are tracked here, all sourced from independent evaluators rather than Meta's own runs, and mostly reported at the model's "xhigh" reasoning-effort setting (as of 2026-08-26): Humanity's Last Exam 45.46% (Artificial Analysis's live model-page data, as of 2026-08-26; AA's separate launch-day article instead states "45% to 44%" as a version-over-version regression versus Muse Spark 1.1, a different, lower figure than the live page's 45.46% rather than a rounding of it — the two AA sources disagree with each other by more than a point), GPQA Diamond 90.4% (Artificial Analysis), Terminal-Bench 2.1 80.15% (Artificial Analysis), SWE-bench Verified 86.6% (vals.ai, rank 13/86, ±1.52 CI, Mini-SWE-agent scaffold — no effort tier is shown on this leaderboard row, so its comparability to the "xhigh" figures elsewhere is unconfirmed), DeepSWE v1.1 55% (deepswe.datacurve.ai, rank 13, ±2% CI), Toolathlon-Verified 75.9% Pass@1 (toolathlon.xyz, rank 3 of 7 rows shown; Pass@3 87.0%, Pass^3 63.0%), and LiveBench 78 (livebench.ai, 2026-06-25 release round; category breakdown: Reasoning 90.0, Mathematics 91.2, Coding 77.5, Data Analysis 76.5, Language 78.6, Instruction Following 74.3, Agentic Coding 57.6).

Terminal-Bench 2.1 and DeepSWE v1.1 both show a vendor-vs-independent spread, but they don't carry equal weight. On Terminal-Bench 2.1, Meta's own launch chart (research.meta.ai) self-reports 82.9% against Artificial Analysis's independent 80.15% — a 2.75-point gap that sits inside this benchmark's own ±10.6 noise band, so on its own it isn't evidence of inflation (vals.ai's independent harness reads lower still, 69.66%, but that's a cross-harness difference between two independent sources, not a vendor-vs-independent one). On DeepSWE v1.1, Meta self-reports 59.3% versus the independent leaderboard's 55% — a +4.3pp gap that DOES exceed that leaderboard's own ±2% confidence interval for this model, making it the page's clearer instance of the self-report/independent divergence this site exists to surface.

Coverage is not complete. GPQA Diamond has not been separately checked on vals.ai's own GPQA Diamond leaderboard — this model is absent from its 20-benchmark coverage list and wasn't found in the top ~28 of its 135 rows. Four other commonly tracked benchmarks were checked live on 2026-08-26 and have no row for this model at all: LiveCodeBench (vals.ai has zero Meta entries for this eval), ARC-AGI-2 (arcprize.org's newest Meta rows are still Llama 4 Scout/Maverick from April 2025), Agents' Last Exam (absent from all 39 rows of its leaderboard), and MathArena's HMMT Feb 2026 (absent from its ~31-model table). Separately, Artificial Analysis flags the model as comparatively verbose — 95M output tokens on its Intelligence Index versus a 72M median for similarly priced models, which matters directly for output-token cost under either pricing tier. Its overall AA Intelligence Index score is 57, ranked #17 of 187 tracked models (artificialanalysis.ai, as of 2026-08-26).

Licensing is a live, unresolved story. Muse Spark 1.2 shipped closed/proprietary on 2026-08-05. On 2026-08-10, Meta Chief AI Officer Alexandr Wang said open weights for Muse Spark 1.2 were "coming soon"; that same day Meta open-weighted a separate, smaller model, Muse Glimmer (30B, Apache 2.0) — not Muse Spark 1.2 itself. As of 2026-08-26, Muse Spark 1.2's weights remain unreleased: both Artificial Analysis ("isOpenWeights": false) and vals.ai ("WEIGHTS: PRIVATE") confirm proprietary status live. This site tracks that as an announced-but-unfulfilled openness promise, not settled fact.

Compare with

FAQ

Is Muse Spark 1.2 good at coding?

Coding scores vary sharply depending on who ran the test, not just on the model itself. On Terminal-Bench 2.1, Meta's own launch materials, Artificial Analysis's independent run, and vals.ai's own harness each report a different score for what is nominally the same benchmark — a 13-point spread from the lowest reading to the highest. Independent evaluators also score it well on SWE-bench Verified and DeepSWE, but this site treats any single coding number for Muse Spark 1.2 as harness-dependent rather than settled; see the tracked scores above for the full breakdown by source.

How does Muse Spark 1.2 perform on reasoning benchmarks like GPQA Diamond and Humanity's Last Exam?

Independent testing from Artificial Analysis puts Muse Spark 1.2 at 90.4% on GPQA Diamond, measured at the model's "xhigh" reasoning-effort setting. It also has a published Humanity's Last Exam score of 45.46% from the same source's live model-page data — a separate, earlier AA article instead states "45% to 44%" for a version comparison against Muse Spark 1.1, which is a different, lower figure rather than a rounding of the live number, so the two AA sources disagree with each other by slightly more than a point. GPQA Diamond doesn't have a second independent evaluator to cross-check against — it's absent from vals.ai's own GPQA Diamond leaderboard for this model — so treat that reading as single-source for now. (HLE's single-source status wasn't separately checked against other leaderboards.)

Is Muse Spark 1.2 open source?

No — Muse Spark 1.2 launched proprietary on 2026-08-05, and it still is as of 2026-08-26. Meta's Chief AI Officer Alexandr Wang said on 2026-08-10 that open weights were "coming soon," and Meta did open-weight a different, smaller model that same day (Muse Glimmer, Apache 2.0) — but not Muse Spark 1.2. As of 2026-08-26, both Artificial Analysis and vals.ai still list it as closed-weight, so treat the open-weighting promise as announced, not delivered.

How much does Muse Spark 1.2 cost to use via the API?

It depends which tier you pick, and the two aren't equivalent. The Standard tier costs $1.25 per 1M input tokens and $4.25 per 1M output tokens, and it carries a no-training-use guarantee. A separate Contributor tier is priced far lower, but it requires opting your prompts and completions into Meta's training data — that's a privacy trade, not just a cheaper plan. See the pricing details tracked on this page for the exact Contributor-tier rates.

What is Muse Code?

Muse Code is a terminal-based coding agent from Meta, released in beta alongside Muse Spark 1.2 on 2026-08-05 and available through dev.meta.ai. It's the delivery vehicle Meta is pairing with this model's coding-focused benchmark push (SWE-bench Verified, Terminal-Bench 2.1, DeepSWE, and Toolathlon-Verified are all tracked on this page), though this site does not yet have independent data on Muse Code's own performance separate from the underlying model.

What is Muse Spark 1.2's best score on a benchmark that still ranks?

Terminal-Bench 2.1 at 80.15 (2026-08-26), with LiveBench at 78 and Toolathlon-Verified at 75.9 behind it. Its two headline numbers, GPQA Diamond 90.4 and SWE-bench Verified 86.6, both come from retired boards here.

Further reading

Benchmark guides