Claude Sonnet 5 logo

Claude Sonnet 5

Anthropic · Released Jun 2026

Recommended

Anthropic's most agentic Sonnet delivers near-Opus 4.8 performance at roughly half the price, with Terminal-Bench 2.1 +13.4 points over Sonnet 4.6. The new tokenizer inflates counts ~30%, so intro pricing ($2/$10 per MTok through Aug 31) keeps migration cost-neutral, but teams should recount budgets before switching. It replaces Sonnet 4.6 as the default across all Claude plans, ideal for agentic coding, brownfield debugging, and production pipelines — not for the hardest frontier reasoning or cybersecurity tasks.

Is it right for you?

Good for

  • Multi-step agentic coding and autonomous task completion — finishes complex workflows previous Sonnet models stalled on
  • Brownfield debugging: traces failures to root causes, writes reproducing tests, and ships durable fixes unprompted
  • Production AI agents at attractive cost — near-Opus 4.8 quality on agentic benchmarks at roughly half the price
  • Knowledge work where it edges Opus 4.8 (GDPval-AA v2: 1,618 vs 1,615) — document drafting, research, professional analysis
  • Self-verification: checks its own output without being explicitly prompted, reducing compounding errors in automated pipelines

Not good for

  • Maximum-capability frontier tasks where Opus 4.8 or Fable 5 are necessary — Sonnet 5 gets close but doesn't match the top tier
  • Cybersecurity work requiring reduced guardrails — cyber safeguards are enabled by default and Anthropic recommends Opus 4.8 for that use case
  • Workloads where the new tokenizer's ~30% token inflation matters — recount budgets before switching; same text costs more tokens than on Sonnet 4.6

How it performs by task

Agentic coding

Excellent

Terminal-Bench 2.1 80.4% (+13.4 pts over Sonnet 4.6) — the clearest signal of multi-step agentic reliability. Cursor production benchmark: 61.2% (vs 49.0% for 4.6).

Code generation

Very Good

SWE-bench Verified 85.2% — 4.1 pts ahead of Sonnet 4.6 (81.1%), closing the gap with Opus 4.8 (88.6%). Strong on standard coding but Opus still leads Pro by 6 pts.

Debugging

Excellent

Early testers repeatedly highlight unprompted reproducing tests, stashing to confirm regressions, and root-cause tracing on brownfield code.

Tool use and computer use

Excellent

OSWorld-Verified 81.2%, matching Opus 4.8 (83.4%) and significantly ahead of Sonnet 4.6 (78.5%). Handles browsers, terminals, and API tools end-to-end.

Knowledge work

Excellent

GDPval-AA v2: 1,618 — slightly edges Opus 4.8 (1,615). A genuine first for a mid-tier Sonnet model against its flagship sibling.

Reasoning

Very Good

Humanity's Last Exam: 43.2% (no tools) / 57.4% (with tools). With tools, nearly matches Opus 4.8 (57.9%); without tools, still trails the flagship tier.

Cybersecurity

Fair

By design — never developed a full working exploit on Firefox 147 vulnerability tests (0.0%). Cyber safeguards enabled by default. Anthropic recommends Opus 4.8 for cyber work.

Pricing

Input

$3 / 1M

Output

$15 / 1M

Context

1M tokens

View full pricing

Benchmarks

BenchmarkScoreSource
SWE-bench Verified85.2% Source
Terminal-Bench 2.180.4% Source
SWE-bench Pro63.2% Source
Humanity's Last Exam (no tools)43.2% Source
Humanity's Last Exam (with tools)57.4% Source
OSWorld-Verified81.2% Source
Knowledge Work (GDPval-AA v2)1,618 Source
Cursor Production Benchmark (Max effort)61.2% Source
BrowseComp 25 (agentic search)84.7% Source

Verdict history

Jul 15, 2026
copy edit to brevity standard, detail preserved in children — EN verdict_summary was ~116 words / 5 sentences — over the 2–4 sentence / ≤90 word target. Shortened to 3 sentences / ~70 words while preserving all claims. Children (strengths, task_strengths, etc.) retain the full detail.

Verification log

  • Pricing— No changes

    Automated agent

  • Pricing— No changes

    Automated agent

    Live $2/$10 intro through Aug 31; stored $3/$15 standard rate. No change to standard rate.

  • Pricing— No changes

    Automated agent

  • Profile— No changes

    Imported at launch

  • Pricing— No changes

    Imported at launch

How we evaluate