Muse Spark 1.2 logo

Muse Spark 1.2

Meta · Released Aug 2026

Conditional

Meta's coding-focused model update, co-trained with the Muse Code harness — 82.9% on Terminal-Bench 2.1 and 59.3% on DeepSWE 1.1, second only to Claude Opus 5 on the former and trailing both Opus 5 and GPT-5.6 Terra on the latter. Standard pricing unchanged from 1.1 at $1.25 input / $4.25 output per 1M tokens; contributor tier drops to $0.10/$0.20 in exchange for training-data rights. All published benchmarks remain vendor-run with no independent reproduction. Best for cost-sensitive coding teams willing to test a model explicitly optimized for its own harness.

Is it right for you?

Good for

  • Cost-sensitive coding agent workloads — $1.25/$4.25 standard pricing undercuts Claude, GPT, and Gemini flagships by 3-5x
  • Muse Code agent integration — co-trained with the harness for best-in-tool performance
  • Long-horizon autonomous coding — 1,000+ tool calls over 24 hours, event-log crash recovery, sustained optimization past initial exploration
  • Multimodal repo understanding — accepts text, image, video, and PDF inputs for coding context

Not good for

  • Production coding accuracy without independent validation — all published benchmarks are vendor-run in Meta's own framework; no third-party reproduction exists yet
  • Non-coding or general agentic tasks — coding-focused checkpoint; reasoning, knowledge, math, and multilingual benchmarks remain unpublished
  • Open-weight or self-hosting requirements — closed weights, API-only; no Hugging Face weights, no fine-tuning
  • High-concurrency production on contributor tier — 60 RPM cap vs 3,000 standard; training-data grant on all contributor traffic

How it performs by task

Code generation

Good

Second on Terminal-Bench 2.1 (3.8 pts behind Opus 5) and Meta's internal coding bench (8.8 pts behind); third on DeepSWE 1.1 behind Opus 5 and GPT-5.6 Terra

Agentic coding

Very Good

Persistent background agents + worktree isolation; 24-hour kernel optimization demo with 1,000+ tool calls; co-trained with Muse Code harness

Long-horizon tasks

Very Good

Sustained improvement over 24 hours on GPU kernel optimization; event-log replay on crash prevents lost work and re-prompting

General reasoning

Fair

Coding-focused checkpoint; reasoning, knowledge, math, and multilingual benchmark categories remain unpublished — cannot assess non-coding strength

Pricing

Input

$1.25 / 1M tokens

Output

$4.25 / 1M tokens

Context

1M tokens

View full pricing

Benchmarks

BenchmarkScoreSource
Terminal-Bench 2.182.9% Source
DeepSWE 1.159.3% Source
Meta Internal Coding Bench70.6% Source
BenchLM Aggregate60.3/100 (#49 of 216) Source

No verdict changes yet

The clock starts day one — changes land here as our verdict evolves.

Verification log

No verification checks yet

We haven't logged a verification check for this entry. Once a check runs, its history shows here.

How we evaluate