Muse Spark 1.2 logo

Muse Spark 1.2

Meta · Sorti août 2026

Conditionnel

La mise à jour du modèle de Meta axée sur le codage, co-entraîné avec le harnais Muse Code — 82,9 % sur Terminal-Bench 2.1 et 59,3 % sur DeepSWE 1.1, deuxième seulement derrière Claude Opus 5 sur le premier. Tarification standard inchangée à 1,25 $ en entrée / 4,25 $ en sortie par million de tokens.

Est-ce fait pour vous ?

Recommandé pour

  • Cost-sensitive coding agent workloads: $1.25/$4.25 standard pricing undercuts Claude, GPT, and Gemini flagships by 3-5x
  • Muse Code agent integration: co-trained with the harness for best-in-tool performance
  • Long-horizon autonomous coding: 1,000+ tool calls over 24 hours, event-log crash recovery, sustained optimization past initial exploration
  • Multimodal repo understanding: accepts text, image, video, and PDF inputs for coding context

Déconseillé pour

  • Production coding accuracy without independent validation: all published benchmarks are vendor-run in Meta's own framework; no third-party reproduction exists yet
  • Non-coding or general agentic tasks: coding-focused checkpoint; reasoning, knowledge, math, and multilingual benchmarks remain unpublished
  • Open-weight or self-hosting requirements: closed weights, API-only; no Hugging Face weights, no fine-tuning
  • High-concurrency production on contributor tier: 100 RPM cap vs 3,000 standard per Meta's rate-limit docs; training-data grant on all contributor traffic

Performances par tâche

Génération de code

Good

Deuxième sur Terminal-Bench 2.1 (3,8 pts derrière Opus 5) ; troisième sur DeepSWE 1.1

Codage agentic

Very Good

Agents d'arrière-plan persistants + isolation worktree ; co-entraîné avec le harnais Muse Code

Tâches longue durée

Very Good

Amélioration soutenue sur 24 heures en optimisation de noyau GPU

Raisonnement général

Fair

Checkpoint axé sur le codage ; les benchmarks de raisonnement, connaissance et mathématiques restent non publiés

Tarifs

Entrée

$1.25 / 1M tokens

Sortie

$4.25 / 1M tokens

Contexte

1M tokens

Voir tous les tarifs

Tests de performance

TestScoreSource
Terminal-Bench 2.182.9% Source
DeepSWE 1.159.3% Source
Meta Internal Coding Bench70.6% Source
BenchLM Aggregate60.3/100 (#49 of 216) Source

Aucun changement de verdict pour l’instant

L’horloge tourne dès le premier jour : les changements apparaîtront ici à mesure que notre verdict évolue.

Journal de vérification

Aucune vérification pour le moment

Nous n’avons pas encore enregistré de vérification pour cette entrée. Dès qu’une vérification s’exécute, son historique apparaît ici.

Comment nous évaluons