Muse Spark 1.2 logo

Muse Spark 1.2

Meta · Lançado 08/2026

Condicional

Meta's coding-focused model update, co-trained with the Muse Code harness — 82.9% on Terminal-Bench 2.1 and 59.3% on DeepSWE 1.1, second only to Claude Opus 5 on the former and trailing both Opus 5 and GPT-5.6 Terra on the latter. Standard pricing unchanged from 1.1 at $1.25 input / $4.25 output per 1M tokens; contributor tier drops to $0.10/$0.20 in exchange for training-data rights. All published benchmarks remain vendor-run with no independent reproduction. Best for cost-sensitive coding teams willing to test a model explicitly optimized for its own harness.

É adequada para ti?

Bom para

  • Cost-sensitive coding agent workloads — $1.25/$4.25 standard pricing undercuts Claude, GPT, and Gemini flagships by 3-5x
  • Muse Code agent integration — co-trained with the harness for best-in-tool performance
  • Long-horizon autonomous coding — 1,000+ tool calls over 24 hours, event-log crash recovery, sustained optimization past initial exploration
  • Multimodal repo understanding — accepts text, image, video, and PDF inputs for coding context

Não recomendado para

  • Production coding accuracy without independent validation — all published benchmarks are vendor-run in Meta's own framework; no third-party reproduction exists yet
  • Non-coding or general agentic tasks — coding-focused checkpoint; reasoning, knowledge, math, and multilingual benchmarks remain unpublished
  • Open-weight or self-hosting requirements — closed weights, API-only; no Hugging Face weights, no fine-tuning
  • High-concurrency production on contributor tier — 60 RPM cap vs 3,000 standard; training-data grant on all contributor traffic

Desempenho por tarefa

Code generation

Good

Second on Terminal-Bench 2.1 (3.8 pts behind Opus 5) and Meta's internal coding bench (8.8 pts behind); third on DeepSWE 1.1 behind Opus 5 and GPT-5.6 Terra

Agentic coding

Very Good

Persistent background agents + worktree isolation; 24-hour kernel optimization demo with 1,000+ tool calls; co-trained with Muse Code harness

Long-horizon tasks

Very Good

Sustained improvement over 24 hours on GPU kernel optimization; event-log replay on crash prevents lost work and re-prompting

General reasoning

Fair

Coding-focused checkpoint; reasoning, knowledge, math, and multilingual benchmark categories remain unpublished — cannot assess non-coding strength

Preços

Entrada

$1.25 / 1M tokens

Saída

$4.25 / 1M tokens

Contexto

1M tokens

Ver preços completos

Testes de desempenho

TestePontuaçãoFonte
Terminal-Bench 2.182.9% Fonte
DeepSWE 1.159.3% Fonte
Meta Internal Coding Bench70.6% Fonte
BenchLM Aggregate60.3/100 (#49 of 216) Fonte

Ainda sem alterações de veredicto

O relógio começa a contar no primeiro dia: as alterações aparecem aqui à medida que o nosso veredicto evolui.

Registo de verificação

Ainda não há verificações

Ainda não registámos nenhuma verificação para esta entrada. Assim que uma for executada, o seu histórico aparece aqui.

Como avaliamos