Claude Fable 5.1
Anthropic · Released Sep 2026
Anthropic's new frontier flagship is the most capable model publicly available, and it largely fixes the safeguard over-triggering that made Fable 5 frustrating for routine coding. The verdict stays conditional on cost: independent measurement shows more spend per task than Fable 5 despite the cheaper cache reads, because the model thinks longer. Reach for it on demanding reasoning and long-horizon agentic work. For everyday coding, Claude Opus 5 at half the rate remains the better default.
Is it right for you?
Good for
- Demanding reasoning and long-horizon agentic work — Anthropic's own guidance is to escalate here when Opus 5 at higher effort still falls short
- Agentic scientific research: 52.6% on Terminal-Bench-Science 0.1 against Opus 5's 29.0% and Fable 5's 24.7%
- Context-heavy agent loops — cache reads dropped 75% to $0.25 per 1M, which Anthropic puts at up to ~45% cheaper on highly agentic workloads
- Long-document and whole-repo work: 1M-token context with 128K max output
- Security-adjacent coding that Fable 5 refused — cyber safeguards now permit vulnerability discovery, cutting Claude Code interventions ~60% per session
Not good for
- Cost-sensitive everyday coding — Claude Opus 5 is 3 index points behind at $5/$25, half the rate
- Latency-sensitive paths — adaptive thinking is always on and Anthropic rates its comparative latency as the slowest in the lineup
- Budgeting by rate card alone — Artificial Analysis measures ~20% higher cost per task than Fable 5 despite the cache-read cut, because effort drives token count
- Fact-recall work where a confident wrong answer is expensive — at max effort it attempts 93.4% of AA-Omniscience questions, trading more hallucinations for its higher accuracy at no net Omniscience gain
How it performs by task
Agentic coding
55.8% Terminal-Bench 4.0 and 73.4% CursorBench 3.2.0, with far fewer safeguard interruptions than Fable 5
Scientific research agents
52.6% Terminal-Bench-Science 0.1, roughly double Fable 5 and well ahead of Opus 5's 29.0%
Frontier reasoning
Tops the AA Intelligence Index at 66; 65.0% on Humanity's Last Exam with tools
Long-context work
1M-token window, 128K output, and cache reads at $0.25 per 1M make repeated large-context passes affordable
Computer use
41.7% OSWorld 2.0 strict — competent but not the reason to pay this rate
High-volume cheap inference
Wrong tier entirely; Haiku 4.5 at $1/$5 or Sonnet 5 at $2/$10 own this
Pricing
Input
$10 / 1M tokens
Output
$50 / 1M tokens
Context
1M context, 128K output; cache $0.25/1M
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| Artificial Analysis Intelligence Index | 66 | Source |
| Terminal-Bench-Science 0.1 | 52.6% | Source |
| Terminal-Bench 4.0 | 55.8% | Source |
| CursorBench 3.2.0 | 73.4% | Source |
| Humanity's Last Exam (no tools) | 60.9% | Source |
| Humanity's Last Exam (with tools) | 65.0% | Source |
| AutomationBench | 31.4% | Source |
| OSWorld 2.0 (strict) | 41.7% | Source |
| GDPval-AA v2 | 1853 | Source |
| AA-Omniscience accuracy (max effort) | 67.2% | Source |
No verdict changes yet
The clock starts day one — changes land here as our verdict evolves.
Sources
- Anthropic — Claude models overviewSep 2026
- AiCybr — Fable 5.1 / Mythos 5.1 pricing, benchmarks, accessSep 2026
- Anthropic — Claude Fable 5.1 and Claude Mythos 5.1Sep 2026
- Artificial Analysis — Claude Fable 5.1 tops the Intelligence IndexSep 2026
- MarkTechPost — Anthropic releases Claude Fable 5.1 and Mythos 5.1Sep 2026
Verification log
No verification checks yet
We haven't logged a verification check for this entry. Once a check runs, its history shows here.