Muse Spark 1.3 logo

Muse Spark 1.3

Meta · Sorti sept. 2026

Conditionnel

Muse Spark 1.3 is Meta's strongest coding and long-context model yet, and an easy drop-in for anyone already on 1.2: the standard token price is unchanged while software-engineering and million-token retrieval results move sharply up. The deciding caveat is that the headline agent numbers come from a max reasoning configuration Meta has not released, while the shipping tier runs slower and costs more per task. Right for cost-sensitive coding teams inside Muse Code; wrong for anyone who needs independently reproduced benchmarks.

Est-ce fait pour vous ?

Recommandé pour

  • Million-token retrieval — MRCR v2 98.5 (256K-512K) and 98.1 (512K-1M) on Meta's launch scorecard, up from 66.3 and 55.5 for Muse Spark 1.2
  • Codebase-scale software engineering — DeepSWE v1.1 75.4, up from 55.0 for Muse Spark 1.2 per Kingy's tally, and SWE-Atlas Codebase Q&A 59.4 (+13.2pp). Peer figures for GPT-5.6 Sol and Claude Opus 5 come from their owners' leaderboards under different harnesses, so the cross-model ranking is not like-for-like
  • Efficiency-sensitive agent runs in Muse Code — Meta reports ~20% fewer tool calls and ~25% fewer tokens than 1.2 on its engineering comparisons
  • Teams that can trade data rights for price — the muse-spark-1.3-contributor tier prices at $0.10 input / $0.20 output per 1M in exchange for training rights on prompts and completions
  • Multimodal repository context — text, image, video and PDF inputs at a 1,048,576-token context window per Meta's model docs

Déconseillé pour

  • Budgeting on per-token rates alone — Artificial Analysis measured cost per task rising from $0.40 on 1.2 to $0.55 on 1.3 despite unchanged rates, from heavier input-token use on agentic evaluations
  • Anyone buying the launch chart's agent results today — the strongest agent scores (OSWorld 2.0 66.9, GDPval-AA v2 1,754 Elo) come from a max reasoning configuration still in safety testing with no API provider access
  • High-concurrency production on the contributor tier — 100 RPM versus 3,000 RPM on standard, per Meta's rate-limit docs
  • Buyers who require independent reproduction of the coding and long-context claims — every headline benchmark is Meta-run; only Artificial Analysis's index and cost-per-task figures are independently measured
  • Unattended xhigh runs on a fixed token budget — Kingy's review found xhigh exhausting budgets without returning usable code artifacts, while high reasoning accepted 8 of 9 tasks
  • Audio workloads — Meta's own docs state audio understanding is not fully supported in 1.3 and requests containing audio may return degraded quality
  • Latency-sensitive interactive work — Kingy measured mean task wall time of 99.0 seconds at xhigh versus 36.0 seconds for Muse Spark 1.2

Performances par tâche

Software engineering (DeepSWE v1.1)

Excellent

75.4 pass rate against Claude Opus 5's 74.0 and GPT-5.6 Sol's 73.0 as collated on Meta's scorecard; the peer figures come from their owners' leaderboards under different harnesses, and the source's ranking prose conflicts with its own numbers. Up from 55.0 for Muse Spark 1.2 per Kingy's tally

Terminal/CLI agent work (Terminal-Bench 2.1)

Very Good

88.8 pass@1, tied with GPT-5.6 Sol and ahead of Claude Opus 5 (86.7) on Meta's chart

Codebase question answering (SWE-Atlas)

Very Good

59.4 pass@1 versus 53.5 for Sol and 52.7 for Opus 5; +13.2pp over 1.2

Long-context retrieval (MRCR v2)

Excellent

98.5 at 256K-512K and 98.1 at 512K-1M — the largest generational jump in the release, from 66.3 and 55.5

Computer use (OSWorld 2.0)

Good

57.2 on the shipping xhigh configuration; the 66.9 figure quoted around the launch belongs to the unreleased max configuration

General intelligence (Artificial Analysis Index v4.1.1)

Good

61 at xhigh on Artificial Analysis's independent measurement, not a frontier-leading score

Audio understanding

Poor

Explicitly not fully supported in 1.3 per Meta's model documentation

Tarifs

Entrée

$1.25 / 1M tokens

Sortie

$4.25 / 1M tokens

Contexte

1M ctx; contributor tier $0.10/$0.20

Voir tous les tarifs

Tests de performance

TestScoreSource
Terminal-Bench 2.188.8 Source
DeepSWE v1.175.4 Source
SWE-Atlas Codebase Q&A59.4 Source
MRCR v2 (256K-512K)98.5 Source
MRCR v2 (512K-1M)98.1 Source
Artificial Analysis Intelligence Index v4.1.1 (xhigh)61 Source
OSWorld 2.0 (shipping xhigh)57.2 Source
JobBench (shipping xhigh)61.2 Source
GDPval-AA v2 (shipping xhigh)1709 Elo Source
Cost per task (Artificial Analysis)$0.55 Source

Aucun changement de verdict pour l’instant

L’horloge tourne dès le premier jour : les changements apparaîtront ici à mesure que notre verdict évolue.

Journal de vérification

Aucune vérification pour le moment

Nous n’avons pas encore enregistré de vérification pour cette entrée. Dès qu’une vérification s’exécute, son historique apparaît ici.

Comment nous évaluons