Muse Spark 1.3 logo

Muse Spark 1.3

Meta · Lanzado sept 2026

Condicional

Muse Spark 1.3 is Meta's strongest coding and long-context model yet, and an easy drop-in for anyone already on 1.2: the standard token price is unchanged while software-engineering and million-token retrieval results move sharply up. The deciding caveat is that the headline agent numbers come from a max reasoning configuration Meta has not released, while the shipping tier runs slower and costs more per task. Right for cost-sensitive coding teams inside Muse Code; wrong for anyone who needs independently reproduced benchmarks.

¿Es adecuada para ti?

Bueno para

  • Million-token retrieval — MRCR v2 98.5 (256K-512K) and 98.1 (512K-1M) on Meta's launch scorecard, up from 66.3 and 55.5 for Muse Spark 1.2
  • Codebase-scale software engineering — DeepSWE v1.1 75.4, up from 55.0 for Muse Spark 1.2 per Kingy's tally, and SWE-Atlas Codebase Q&A 59.4 (+13.2pp). Peer figures for GPT-5.6 Sol and Claude Opus 5 come from their owners' leaderboards under different harnesses, so the cross-model ranking is not like-for-like
  • Efficiency-sensitive agent runs in Muse Code — Meta reports ~20% fewer tool calls and ~25% fewer tokens than 1.2 on its engineering comparisons
  • Teams that can trade data rights for price — the muse-spark-1.3-contributor tier prices at $0.10 input / $0.20 output per 1M in exchange for training rights on prompts and completions
  • Multimodal repository context — text, image, video and PDF inputs at a 1,048,576-token context window per Meta's model docs

No recomendado para

  • Budgeting on per-token rates alone — Artificial Analysis measured cost per task rising from $0.40 on 1.2 to $0.55 on 1.3 despite unchanged rates, from heavier input-token use on agentic evaluations
  • Anyone buying the launch chart's agent results today — the strongest agent scores (OSWorld 2.0 66.9, GDPval-AA v2 1,754 Elo) come from a max reasoning configuration still in safety testing with no API provider access
  • High-concurrency production on the contributor tier — 100 RPM versus 3,000 RPM on standard, per Meta's rate-limit docs
  • Buyers who require independent reproduction of the coding and long-context claims — every headline benchmark is Meta-run; only Artificial Analysis's index and cost-per-task figures are independently measured
  • Unattended xhigh runs on a fixed token budget — Kingy's review found xhigh exhausting budgets without returning usable code artifacts, while high reasoning accepted 8 of 9 tasks
  • Audio workloads — Meta's own docs state audio understanding is not fully supported in 1.3 and requests containing audio may return degraded quality
  • Latency-sensitive interactive work — Kingy measured mean task wall time of 99.0 seconds at xhigh versus 36.0 seconds for Muse Spark 1.2

Rendimiento por tarea

Software engineering (DeepSWE v1.1)

Excellent

75.4 pass rate against Claude Opus 5's 74.0 and GPT-5.6 Sol's 73.0 as collated on Meta's scorecard; the peer figures come from their owners' leaderboards under different harnesses, and the source's ranking prose conflicts with its own numbers. Up from 55.0 for Muse Spark 1.2 per Kingy's tally

Terminal/CLI agent work (Terminal-Bench 2.1)

Very Good

88.8 pass@1, tied with GPT-5.6 Sol and ahead of Claude Opus 5 (86.7) on Meta's chart

Codebase question answering (SWE-Atlas)

Very Good

59.4 pass@1 versus 53.5 for Sol and 52.7 for Opus 5; +13.2pp over 1.2

Long-context retrieval (MRCR v2)

Excellent

98.5 at 256K-512K and 98.1 at 512K-1M — the largest generational jump in the release, from 66.3 and 55.5

Computer use (OSWorld 2.0)

Good

57.2 on the shipping xhigh configuration; the 66.9 figure quoted around the launch belongs to the unreleased max configuration

General intelligence (Artificial Analysis Index v4.1.1)

Good

61 at xhigh on Artificial Analysis's independent measurement, not a frontier-leading score

Audio understanding

Poor

Explicitly not fully supported in 1.3 per Meta's model documentation

Precios

Entrada

$1.25 / 1M tokens

Salida

$4.25 / 1M tokens

Contexto

1M ctx; contributor tier $0.10/$0.20

Ver precios completos

Pruebas de rendimiento

PruebaPuntuaciónFuente
Terminal-Bench 2.188.8 Fuente
DeepSWE v1.175.4 Fuente
SWE-Atlas Codebase Q&A59.4 Fuente
MRCR v2 (256K-512K)98.5 Fuente
MRCR v2 (512K-1M)98.1 Fuente
Artificial Analysis Intelligence Index v4.1.1 (xhigh)61 Fuente
OSWorld 2.0 (shipping xhigh)57.2 Fuente
JobBench (shipping xhigh)61.2 Fuente
GDPval-AA v2 (shipping xhigh)1709 Elo Fuente
Cost per task (Artificial Analysis)$0.55 Fuente

Aún no hay cambios de veredicto

El reloj corre desde el primer día: los cambios aparecerán aquí a medida que evolucione nuestro veredicto.

Registro de verificación

Aún no hay comprobaciones de verificación

Todavía no hemos registrado ninguna comprobación de verificación para esta entrada. Cuando se ejecute una, su historial aparecerá aquí.

Cómo evaluamos