Claude Sonnet 5.5 logo

Claude Sonnet 5.5

Anthropic · Released Sep 2026

Conditional

Conditional. Claude Sonnet 5.5 is Anthropic's default Sonnet model on the Anthropic API, on the same rate card as Sonnet 5, and Anthropic says it runs faster and costs less per task. Every benchmark is Anthropic's own, and five breaking changes hit code written for Sonnet 5. Right for teams ready to re-test agentic coding and knowledge work; hold off if you rely on forced tool choice or temperature control.

Is it right for you?

Good for

  • Agentic coding and command-line work; Anthropic reports 70.6% on Terminal-Bench 4.0, unreproduced by us
  • Large-context work: 1M-token context and 128K output, 300K output on the Batches API with a beta header
  • Cost control at scale: $2 input and $10 output per 1M tokens, batch at half price and cache reads at $0.20
  • Multi-cloud deployment: the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS
  • Claude Code users: Sonnet 5.5 (claude-sonnet-5-5) was added in Claude Code 2.1.284 on Sep 28, 2026

Not good for

  • Code that forces tool use: forced tool use returns an error
  • Apps that set temperature, top_p or top_k: non-default values return a 400 error
  • Interfaces that stream text between tool calls without setting a display value or between_tools: they go quiet
  • Higher-risk cybersecurity tasks, which Anthropic says visibly fall back to Sonnet 5
  • Anything beyond text and image input with text output

How it performs by task

Agentic coding

Very Good

Anthropic-reported Terminal-Bench 4.0 and CursorBench 4.0 results; vendor-reported, untested by us

Knowledge work

Very Good

Anthropic-reported GDPval-AA v2.1 of 1844; vendor-reported, untested by us

Cybersecurity

Fair

Higher-risk cyber tasks visibly fall back to Sonnet 5 by design

Pricing

Input

$2 / 1M tokens

Output

$10 / 1M tokens

Context

1M ctx; batch $1/$5; cache read $0.20

View full pricing

Benchmarks

BenchmarkScoreSource
Terminal-Bench 4.0 (Anthropic-reported)70.6% Source
CursorBench 4.0 (Anthropic-reported)55.5% Source
GDPval-AA v2.1 (Anthropic-reported)1844 Source
Humanity's Last Exam, with tools (Anthropic-reported)64.5% Source

No verdict changes yet

The clock starts day one. Changes land here as our verdict evolves.

Verification log

  • Pricing· No changes

    Automated agent

  • Pricing· No changes

    Automated agent

How we evaluate