Claude Fable 5.1 logo

Claude Fable 5.1

Anthropic · Released Sep 2026

Conditional

Anthropic's new frontier flagship is the most capable model publicly available, and it largely fixes the safeguard over-triggering that made Fable 5 frustrating for routine coding. The verdict stays conditional on cost: independent measurement shows more spend per task than Fable 5 despite the cheaper cache reads, because the model thinks longer. Reach for it on demanding reasoning and long-horizon agentic work. For everyday coding, Claude Opus 5 at half the rate remains the better default.

Is it right for you?

Good for

  • Demanding reasoning and long-horizon agentic work — Anthropic's own guidance is to escalate here when Opus 5 at higher effort still falls short
  • Agentic scientific research: 52.6% on Terminal-Bench-Science 0.1 against Opus 5's 29.0% and Fable 5's 24.7%
  • Context-heavy agent loops — cache reads dropped 75% to $0.25 per 1M, which Anthropic puts at up to ~45% cheaper on highly agentic workloads
  • Long-document and whole-repo work: 1M-token context with 128K max output
  • Security-adjacent coding that Fable 5 refused — cyber safeguards now permit vulnerability discovery, cutting Claude Code interventions ~60% per session

Not good for

  • Cost-sensitive everyday coding — Claude Opus 5 is 3 index points behind at $5/$25, half the rate
  • Latency-sensitive paths — adaptive thinking is always on and Anthropic rates its comparative latency as the slowest in the lineup
  • Budgeting by rate card alone — Artificial Analysis measures ~20% higher cost per task than Fable 5 despite the cache-read cut, because effort drives token count
  • Fact-recall work where a confident wrong answer is expensive — at max effort it attempts 93.4% of AA-Omniscience questions, trading more hallucinations for its higher accuracy at no net Omniscience gain

How it performs by task

Agentic coding

Excellent

55.8% Terminal-Bench 4.0 and 73.4% CursorBench 3.2.0, with far fewer safeguard interruptions than Fable 5

Scientific research agents

Excellent

52.6% Terminal-Bench-Science 0.1, roughly double Fable 5 and well ahead of Opus 5's 29.0%

Frontier reasoning

Excellent

Tops the AA Intelligence Index at 66; 65.0% on Humanity's Last Exam with tools

Long-context work

Excellent

1M-token window, 128K output, and cache reads at $0.25 per 1M make repeated large-context passes affordable

Computer use

Good

41.7% OSWorld 2.0 strict — competent but not the reason to pay this rate

High-volume cheap inference

Poor

Wrong tier entirely; Haiku 4.5 at $1/$5 or Sonnet 5 at $2/$10 own this

Pricing

Input

$10 / 1M tokens

Output

$50 / 1M tokens

Context

1M context, 128K output; cache $0.25/1M

View full pricing

Benchmarks

BenchmarkScoreSource
Artificial Analysis Intelligence Index66 Source
Terminal-Bench-Science 0.152.6% Source
Terminal-Bench 4.055.8% Source
CursorBench 3.2.073.4% Source
Humanity's Last Exam (no tools)60.9% Source
Humanity's Last Exam (with tools)65.0% Source
AutomationBench31.4% Source
OSWorld 2.0 (strict)41.7% Source
GDPval-AA v21853 Source
AA-Omniscience accuracy (max effort)67.2% Source

No verdict changes yet

The clock starts day one — changes land here as our verdict evolves.

Verification log

No verification checks yet

We haven't logged a verification check for this entry. Once a check runs, its history shows here.

How we evaluate