Muse Spark 1.3
Meta · Veröffentlicht Sept. 2026
Muse Spark 1.3 is Meta's strongest coding and long-context model yet, and an easy drop-in for anyone already on 1.2: the standard token price is unchanged while software-engineering and million-token retrieval results move sharply up. The deciding caveat is that the headline agent numbers come from a max reasoning configuration Meta has not released, while the shipping tier runs slower and costs more per task. Right for cost-sensitive coding teams inside Muse Code; wrong for anyone who needs independently reproduced benchmarks.
Ist es das Richtige für dich?
Gut für
- Million-token retrieval — MRCR v2 98.5 (256K-512K) and 98.1 (512K-1M) on Meta's launch scorecard, up from 66.3 and 55.5 for Muse Spark 1.2
- Codebase-scale software engineering — DeepSWE v1.1 75.4, up from 55.0 for Muse Spark 1.2 per Kingy's tally, and SWE-Atlas Codebase Q&A 59.4 (+13.2pp). Peer figures for GPT-5.6 Sol and Claude Opus 5 come from their owners' leaderboards under different harnesses, so the cross-model ranking is not like-for-like
- Efficiency-sensitive agent runs in Muse Code — Meta reports ~20% fewer tool calls and ~25% fewer tokens than 1.2 on its engineering comparisons
- Teams that can trade data rights for price — the muse-spark-1.3-contributor tier prices at $0.10 input / $0.20 output per 1M in exchange for training rights on prompts and completions
- Multimodal repository context — text, image, video and PDF inputs at a 1,048,576-token context window per Meta's model docs
Nicht geeignet für
- Budgeting on per-token rates alone — Artificial Analysis measured cost per task rising from $0.40 on 1.2 to $0.55 on 1.3 despite unchanged rates, from heavier input-token use on agentic evaluations
- Anyone buying the launch chart's agent results today — the strongest agent scores (OSWorld 2.0 66.9, GDPval-AA v2 1,754 Elo) come from a max reasoning configuration still in safety testing with no API provider access
- High-concurrency production on the contributor tier — 100 RPM versus 3,000 RPM on standard, per Meta's rate-limit docs
- Buyers who require independent reproduction of the coding and long-context claims — every headline benchmark is Meta-run; only Artificial Analysis's index and cost-per-task figures are independently measured
- Unattended xhigh runs on a fixed token budget — Kingy's review found xhigh exhausting budgets without returning usable code artifacts, while high reasoning accepted 8 of 9 tasks
- Audio workloads — Meta's own docs state audio understanding is not fully supported in 1.3 and requests containing audio may return degraded quality
- Latency-sensitive interactive work — Kingy measured mean task wall time of 99.0 seconds at xhigh versus 36.0 seconds for Muse Spark 1.2
Leistung nach Aufgabe
Software engineering (DeepSWE v1.1)
75.4 pass rate against Claude Opus 5's 74.0 and GPT-5.6 Sol's 73.0 as collated on Meta's scorecard; the peer figures come from their owners' leaderboards under different harnesses, and the source's ranking prose conflicts with its own numbers. Up from 55.0 for Muse Spark 1.2 per Kingy's tally
Terminal/CLI agent work (Terminal-Bench 2.1)
88.8 pass@1, tied with GPT-5.6 Sol and ahead of Claude Opus 5 (86.7) on Meta's chart
Codebase question answering (SWE-Atlas)
59.4 pass@1 versus 53.5 for Sol and 52.7 for Opus 5; +13.2pp over 1.2
Long-context retrieval (MRCR v2)
98.5 at 256K-512K and 98.1 at 512K-1M — the largest generational jump in the release, from 66.3 and 55.5
Computer use (OSWorld 2.0)
57.2 on the shipping xhigh configuration; the 66.9 figure quoted around the launch belongs to the unreleased max configuration
General intelligence (Artificial Analysis Index v4.1.1)
61 at xhigh on Artificial Analysis's independent measurement, not a frontier-leading score
Audio understanding
Explicitly not fully supported in 1.3 per Meta's model documentation
Preise
Eingabe
$1.25 / 1M tokens
Ausgabe
$4.25 / 1M tokens
Kontext
1M ctx; contributor tier $0.10/$0.20
Benchmarks
| Benchmark | Wert | Quelle |
|---|---|---|
| Terminal-Bench 2.1 | 88.8 | Quelle |
| DeepSWE v1.1 | 75.4 | Quelle |
| SWE-Atlas Codebase Q&A | 59.4 | Quelle |
| MRCR v2 (256K-512K) | 98.5 | Quelle |
| MRCR v2 (512K-1M) | 98.1 | Quelle |
| Artificial Analysis Intelligence Index v4.1.1 (xhigh) | 61 | Quelle |
| OSWorld 2.0 (shipping xhigh) | 57.2 | Quelle |
| JobBench (shipping xhigh) | 61.2 | Quelle |
| GDPval-AA v2 (shipping xhigh) | 1709 Elo | Quelle |
| Cost per task (Artificial Analysis) | $0.55 | Quelle |
Noch keine Urteilsänderungen
Die Uhr läuft ab dem ersten Tag — Änderungen erscheinen hier, sobald sich unser Urteil weiterentwickelt.
Quellen
- Kingy AI — Muse Spark 1.3 review: benchmarks, pricing, verdictSept. 2026
- BenchLM.ai — Muse Spark 1.3 profileSept. 2026
- Meta Model API — models documentation (official)Aug. 2026
- Meta Model API — pricing and rate limits (official)Aug. 2026
- VentureBeat — Meta says Muse Spark 1.3 has frontier performance, but its best results come from a model developers can't broadly use yetSept. 2026
- Meta AI Research — Introducing Muse Spark 1.3 (official announcement)Sept. 2026
Prüfprotokoll
Noch keine Prüfungen
Wir haben für diesen Eintrag noch keine Prüfung erfasst. Sobald eine Prüfung läuft, erscheint ihr Verlauf hier.