Muse Spark 1.3 logo

Muse Spark 1.3

Meta · Veröffentlicht Sept. 2026

Bedingt

Muse Spark 1.3 is Meta's strongest coding and long-context model yet, and an easy drop-in for anyone already on 1.2: the standard token price is unchanged while software-engineering and million-token retrieval results move sharply up. The deciding caveat is that the headline agent numbers come from a max reasoning configuration Meta has not released, while the shipping tier runs slower and costs more per task. Right for cost-sensitive coding teams inside Muse Code; wrong for anyone who needs independently reproduced benchmarks.

Ist es das Richtige für dich?

Gut für

  • Million-token retrieval — MRCR v2 98.5 (256K-512K) and 98.1 (512K-1M) on Meta's launch scorecard, up from 66.3 and 55.5 for Muse Spark 1.2
  • Codebase-scale software engineering — DeepSWE v1.1 75.4, up from 55.0 for Muse Spark 1.2 per Kingy's tally, and SWE-Atlas Codebase Q&A 59.4 (+13.2pp). Peer figures for GPT-5.6 Sol and Claude Opus 5 come from their owners' leaderboards under different harnesses, so the cross-model ranking is not like-for-like
  • Efficiency-sensitive agent runs in Muse Code — Meta reports ~20% fewer tool calls and ~25% fewer tokens than 1.2 on its engineering comparisons
  • Teams that can trade data rights for price — the muse-spark-1.3-contributor tier prices at $0.10 input / $0.20 output per 1M in exchange for training rights on prompts and completions
  • Multimodal repository context — text, image, video and PDF inputs at a 1,048,576-token context window per Meta's model docs

Nicht geeignet für

  • Budgeting on per-token rates alone — Artificial Analysis measured cost per task rising from $0.40 on 1.2 to $0.55 on 1.3 despite unchanged rates, from heavier input-token use on agentic evaluations
  • Anyone buying the launch chart's agent results today — the strongest agent scores (OSWorld 2.0 66.9, GDPval-AA v2 1,754 Elo) come from a max reasoning configuration still in safety testing with no API provider access
  • High-concurrency production on the contributor tier — 100 RPM versus 3,000 RPM on standard, per Meta's rate-limit docs
  • Buyers who require independent reproduction of the coding and long-context claims — every headline benchmark is Meta-run; only Artificial Analysis's index and cost-per-task figures are independently measured
  • Unattended xhigh runs on a fixed token budget — Kingy's review found xhigh exhausting budgets without returning usable code artifacts, while high reasoning accepted 8 of 9 tasks
  • Audio workloads — Meta's own docs state audio understanding is not fully supported in 1.3 and requests containing audio may return degraded quality
  • Latency-sensitive interactive work — Kingy measured mean task wall time of 99.0 seconds at xhigh versus 36.0 seconds for Muse Spark 1.2

Leistung nach Aufgabe

Software engineering (DeepSWE v1.1)

Excellent

75.4 pass rate against Claude Opus 5's 74.0 and GPT-5.6 Sol's 73.0 as collated on Meta's scorecard; the peer figures come from their owners' leaderboards under different harnesses, and the source's ranking prose conflicts with its own numbers. Up from 55.0 for Muse Spark 1.2 per Kingy's tally

Terminal/CLI agent work (Terminal-Bench 2.1)

Very Good

88.8 pass@1, tied with GPT-5.6 Sol and ahead of Claude Opus 5 (86.7) on Meta's chart

Codebase question answering (SWE-Atlas)

Very Good

59.4 pass@1 versus 53.5 for Sol and 52.7 for Opus 5; +13.2pp over 1.2

Long-context retrieval (MRCR v2)

Excellent

98.5 at 256K-512K and 98.1 at 512K-1M — the largest generational jump in the release, from 66.3 and 55.5

Computer use (OSWorld 2.0)

Good

57.2 on the shipping xhigh configuration; the 66.9 figure quoted around the launch belongs to the unreleased max configuration

General intelligence (Artificial Analysis Index v4.1.1)

Good

61 at xhigh on Artificial Analysis's independent measurement, not a frontier-leading score

Audio understanding

Poor

Explicitly not fully supported in 1.3 per Meta's model documentation

Preise

Eingabe

$1.25 / 1M tokens

Ausgabe

$4.25 / 1M tokens

Kontext

1M ctx; contributor tier $0.10/$0.20

Alle Preise ansehen

Benchmarks

BenchmarkWertQuelle
Terminal-Bench 2.188.8 Quelle
DeepSWE v1.175.4 Quelle
SWE-Atlas Codebase Q&A59.4 Quelle
MRCR v2 (256K-512K)98.5 Quelle
MRCR v2 (512K-1M)98.1 Quelle
Artificial Analysis Intelligence Index v4.1.1 (xhigh)61 Quelle
OSWorld 2.0 (shipping xhigh)57.2 Quelle
JobBench (shipping xhigh)61.2 Quelle
GDPval-AA v2 (shipping xhigh)1709 Elo Quelle
Cost per task (Artificial Analysis)$0.55 Quelle

Noch keine Urteilsänderungen

Die Uhr läuft ab dem ersten Tag — Änderungen erscheinen hier, sobald sich unser Urteil weiterentwickelt.

Prüfprotokoll

Noch keine Prüfungen

Wir haben für diesen Eintrag noch keine Prüfung erfasst. Sobald eine Prüfung läuft, erscheint ihr Verlauf hier.

Wie wir bewerten