Muse Spark 1.3
Meta · Lançado 09/2026
Muse Spark 1.3 is Meta's strongest coding and long-context model yet, and an easy drop-in for anyone already on 1.2: the standard token price is unchanged while software-engineering and million-token retrieval results move sharply up. The deciding caveat is that the headline agent numbers come from a max reasoning configuration Meta has not released, while the shipping tier runs slower and costs more per task. Right for cost-sensitive coding teams inside Muse Code; wrong for anyone who needs independently reproduced benchmarks.
É adequada para ti?
Bom para
- Million-token retrieval — MRCR v2 98.5 (256K-512K) and 98.1 (512K-1M) on Meta's launch scorecard, up from 66.3 and 55.5 for Muse Spark 1.2
- Codebase-scale software engineering — DeepSWE v1.1 75.4, up from 55.0 for Muse Spark 1.2 per Kingy's tally, and SWE-Atlas Codebase Q&A 59.4 (+13.2pp). Peer figures for GPT-5.6 Sol and Claude Opus 5 come from their owners' leaderboards under different harnesses, so the cross-model ranking is not like-for-like
- Efficiency-sensitive agent runs in Muse Code — Meta reports ~20% fewer tool calls and ~25% fewer tokens than 1.2 on its engineering comparisons
- Teams that can trade data rights for price — the muse-spark-1.3-contributor tier prices at $0.10 input / $0.20 output per 1M in exchange for training rights on prompts and completions
- Multimodal repository context — text, image, video and PDF inputs at a 1,048,576-token context window per Meta's model docs
Não recomendado para
- Budgeting on per-token rates alone — Artificial Analysis measured cost per task rising from $0.40 on 1.2 to $0.55 on 1.3 despite unchanged rates, from heavier input-token use on agentic evaluations
- Anyone buying the launch chart's agent results today — the strongest agent scores (OSWorld 2.0 66.9, GDPval-AA v2 1,754 Elo) come from a max reasoning configuration still in safety testing with no API provider access
- High-concurrency production on the contributor tier — 100 RPM versus 3,000 RPM on standard, per Meta's rate-limit docs
- Buyers who require independent reproduction of the coding and long-context claims — every headline benchmark is Meta-run; only Artificial Analysis's index and cost-per-task figures are independently measured
- Unattended xhigh runs on a fixed token budget — Kingy's review found xhigh exhausting budgets without returning usable code artifacts, while high reasoning accepted 8 of 9 tasks
- Audio workloads — Meta's own docs state audio understanding is not fully supported in 1.3 and requests containing audio may return degraded quality
- Latency-sensitive interactive work — Kingy measured mean task wall time of 99.0 seconds at xhigh versus 36.0 seconds for Muse Spark 1.2
Desempenho por tarefa
Software engineering (DeepSWE v1.1)
75.4 pass rate against Claude Opus 5's 74.0 and GPT-5.6 Sol's 73.0 as collated on Meta's scorecard; the peer figures come from their owners' leaderboards under different harnesses, and the source's ranking prose conflicts with its own numbers. Up from 55.0 for Muse Spark 1.2 per Kingy's tally
Terminal/CLI agent work (Terminal-Bench 2.1)
88.8 pass@1, tied with GPT-5.6 Sol and ahead of Claude Opus 5 (86.7) on Meta's chart
Codebase question answering (SWE-Atlas)
59.4 pass@1 versus 53.5 for Sol and 52.7 for Opus 5; +13.2pp over 1.2
Long-context retrieval (MRCR v2)
98.5 at 256K-512K and 98.1 at 512K-1M — the largest generational jump in the release, from 66.3 and 55.5
Computer use (OSWorld 2.0)
57.2 on the shipping xhigh configuration; the 66.9 figure quoted around the launch belongs to the unreleased max configuration
General intelligence (Artificial Analysis Index v4.1.1)
61 at xhigh on Artificial Analysis's independent measurement, not a frontier-leading score
Audio understanding
Explicitly not fully supported in 1.3 per Meta's model documentation
Preços
Entrada
$1.25 / 1M tokens
Saída
$4.25 / 1M tokens
Contexto
1M ctx; contributor tier $0.10/$0.20
Testes de desempenho
| Teste | Pontuação | Fonte |
|---|---|---|
| Terminal-Bench 2.1 | 88.8 | Fonte |
| DeepSWE v1.1 | 75.4 | Fonte |
| SWE-Atlas Codebase Q&A | 59.4 | Fonte |
| MRCR v2 (256K-512K) | 98.5 | Fonte |
| MRCR v2 (512K-1M) | 98.1 | Fonte |
| Artificial Analysis Intelligence Index v4.1.1 (xhigh) | 61 | Fonte |
| OSWorld 2.0 (shipping xhigh) | 57.2 | Fonte |
| JobBench (shipping xhigh) | 61.2 | Fonte |
| GDPval-AA v2 (shipping xhigh) | 1709 Elo | Fonte |
| Cost per task (Artificial Analysis) | $0.55 | Fonte |
Ainda sem alterações de veredicto
O relógio começa a contar no primeiro dia: as alterações aparecem aqui à medida que o nosso veredicto evolui.
Fontes
- Kingy AI — Muse Spark 1.3 review: benchmarks, pricing, verdict09/2026
- BenchLM.ai — Muse Spark 1.3 profile09/2026
- Meta Model API — models documentation (official)08/2026
- Meta Model API — pricing and rate limits (official)08/2026
- VentureBeat — Meta says Muse Spark 1.3 has frontier performance, but its best results come from a model developers can't broadly use yet09/2026
- Meta AI Research — Introducing Muse Spark 1.3 (official announcement)09/2026
Registo de verificação
Ainda não há verificações
Ainda não registámos nenhuma verificação para esta entrada. Assim que uma for executada, o seu histórico aparece aqui.