Claude Sonnet 5.5 logo

Claude Sonnet 5.5

Anthropic · Lanzado sept 2026

Condicional

Conditional. Claude Sonnet 5.5 is Anthropic's default Sonnet model on the Anthropic API, on the same rate card as Sonnet 5, and Anthropic says it runs faster and costs less per task. Every benchmark is Anthropic's own, and five breaking changes hit code written for Sonnet 5. Right for teams ready to re-test agentic coding and knowledge work; hold off if you rely on forced tool choice or temperature control.

¿Es adecuada para ti?

Bueno para

  • Agentic coding and command-line work; Anthropic reports 70.6% on Terminal-Bench 4.0, unreproduced by us
  • Large-context work: 1M-token context and 128K output, 300K output on the Batches API with a beta header
  • Cost control at scale: $2 input and $10 output per 1M tokens, batch at half price and cache reads at $0.20
  • Multi-cloud deployment: the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS
  • Claude Code users: Sonnet 5.5 (claude-sonnet-5-5) was added in Claude Code 2.1.284 on Sep 28, 2026

No recomendado para

  • Code that forces tool use: forced tool use returns an error
  • Apps that set temperature, top_p or top_k: non-default values return a 400 error
  • Interfaces that stream text between tool calls without setting a display value or between_tools: they go quiet
  • Higher-risk cybersecurity tasks, which Anthropic says visibly fall back to Sonnet 5
  • Anything beyond text and image input with text output

Rendimiento por tarea

Agentic coding

Very Good

Anthropic-reported Terminal-Bench 4.0 and CursorBench 4.0 results; vendor-reported, untested by us

Knowledge work

Very Good

Anthropic-reported GDPval-AA v2.1 of 1844; vendor-reported, untested by us

Cybersecurity

Fair

Higher-risk cyber tasks visibly fall back to Sonnet 5 by design

Precios

Entrada

$2 / 1M tokens

Salida

$10 / 1M tokens

Contexto

1M ctx; batch $1/$5; cache read $0.20

Ver precios completos

Pruebas de rendimiento

PruebaPuntuaciónFuente
Terminal-Bench 4.0 (Anthropic-reported)70.6% Fuente
CursorBench 4.0 (Anthropic-reported)55.5% Fuente
GDPval-AA v2.1 (Anthropic-reported)1844 Fuente
Humanity's Last Exam, with tools (Anthropic-reported)64.5% Fuente

Aún no hay cambios de veredicto

El reloj corre desde el primer día: los cambios aparecerán aquí a medida que evolucione nuestro veredicto.

Registro de verificación

  • Precios· Sin cambios

    Agente automatizado

  • Precios· Sin cambios

    Agente automatizado

Cómo evaluamos