Muse Spark 1.1 logo

Muse Spark 1.1

Meta · Released Jul 2026

Conditional

Meta's first paid AI model and first salvo in the commercial coding market — $1.25/$4.25 per 1M tokens with 1M context, native subagent orchestration, MCP/tool support, and computer-use capability at roughly 4x under Anthropic/OpenAI pricing. Leads on tool-use benchmarks but trails Opus 4.8 on pure coding by 7 points on SWE-Bench Pro. Closed-weight, US-only public preview. Best for teams prioritizing tool-use orchestration at low cost.

Is it right for you?

Good for

  • Scaled tool use and MCP orchestration — leads MCP Atlas (88.1) and JobBench (54.7), ahead of Opus 4.8 and GPT-5.5
  • Cost-sensitive agentic workloads — $1.25/$4.25 pricing undercuts Anthropic/OpenAI flagship by ~4x
  • Multi-agent pipelines — native primary-agent/subagent orchestration with parallel execution
  • Multimodal reasoning — handles images, video, documents alongside code; 1M context actively managed

Not good for

  • Frontier coding accuracy — trails Opus 4.8 by 7 pts on SWE-Bench Pro, third on DeepSWE 1.1
  • EU-based or non-US developers — API is US-only public preview; no EU access yet
  • Open-weight or self-hosting requirements — proprietary and closed-weight (unlike Llama)

How it performs by task

Code generation

Good

Trails Opus 4.8 by 7 pts on SWE-Bench Pro; third on DeepSWE 1.1 behind GPT-5.5 and Opus 4.8

Agentic coding

Very Good

Best-in-class tool use (MCP Atlas 88.1, JobBench 54.7); native subagent orchestration with parallel execution

Computer use

Very Good

OSWorld-Verified 80.8 trails Opus 4.8 (83.4); can drive a real desktop using hybrid scripting/clicking approach

Multimodal understanding

Good

Handles images, video, documents; ties Opus 4.8 on several vision benchmarks per vendor charts

Reasoning

Good

SOTA on Humanity's Last Exam and FinanceBench; lags on pure coding reasoning benchmarks

Pricing

Input

$1.25 / 1M tokens

Output

$4.25 / 1M tokens

Context

1M tokens

View full pricing

Benchmarks

BenchmarkScoreSource
MCP Atlas (scaled tool use)88.1 Source
JobBench (professional tool use)54.7 Source
Terminal-Bench 2.180.0 Source
OSWorld-Verified80.8 Source
Humanity's Last ExamSOTA Source

Verdict history

Jul 15, 2026
copy edit to brevity standard — detail preserved in existing structured children

Verification log

  • Pricing— No changes

    Automated agent

  • Pricing— No changes

    Automated agent

  • Profile— No changes

    Imported at launch

  • Pricing— No changes

    Imported at launch

How we evaluate