GPT-5.6 Sol Beats Fable 5 on DeepSWE at One-Third the Cost
- What happened
- GPT-5.6 Sol scored ~72–73% on the DeepSWE coding-agent benchmark to Fable 5's ~70%, at roughly one-third the cost per task ($8.40 vs $13–22).
- Why it matters
- This is the first credible third-party benchmark since both models reached GA, and it reverses the expected outcome: the cheaper model wins on both performance and cost. Fable 5's $10/$50 credit-only pricing (starting tonight) is now empirically harder to justify.
- What to do
- If you're choosing between GPT-5.6 Sol and Fable 5 for coding-agent workloads, the benchmark and pricing data both favor Sol. Run your own representative tasks — but the directional signal is clear.
GPT-5.6 Sol has beaten Claude Fable 5 on DeepSWE — the most rigorous coding-agent benchmark available — while costing roughly one-third as much per task. The numbers arrive just as Fable 5's subscription-included access window closes tonight.
DeepSWE is not a code-generation contest. It drops an agent into a real open-source repository with a requested change and verifies whether the final patch actually works — 113 tasks across 91 repositories and 5 languages, testing repo navigation, bug diagnosis, code editing, tool use, and multi-step repair. The score means "what percentage of real software engineering tasks were solved." GPT-5.6 Sol solved more of them, for less money.
What happened: GPT-5.6 Sol tops DeepSWE benchmark at one-third the cost
Independent evaluator Rohan Paul published the first credible third-party DeepSWE benchmark comparison since both models reached general availability. GPT-5.6 Sol (max) scored approximately 72–73% at roughly $8.40 per task, while Claude Fable 5 (max) scored approximately 70% at $13–22 per task (Rohan Paul(opens in new tab), 2026).
| Model | DeepSWE score | Cost per task | Input / Output pricing |
|---|---|---|---|
| GPT-5.6 Sol (max) | ~72–73% | ~$8.40 | $5 / $30 per 1M tokens |
| Claude Fable 5 (max) | ~70% | $13–22 | $10 / $50 per 1M tokens |
The Artificial Analysis Coding Agent Index — which combines DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA across model-specific harnesses — corroborates the finding: GPT-5.6 Sol (max) in Codex leads with 80 points, ahead of Claude Fable 5 (max) in Claude Code at 77 (Artificial Analysis(opens in new tab), 2026).
On the broader Intelligence Index, Sol scores 59 to Fable 5's 60 — one point behind at one-third the cost per task ($1.04 vs $2.75).
The cost gap across Anthropic's lineup is even starker:
- GPT-5.6 Sol: 3x cheaper than Fable 5
- GPT-5.6 Sol: 5–7x cheaper than Opus 4.8
- GPT-5.6 Sol: 10x cheaper than Sonnet 5
These multipliers come from the International Cyber Digest(opens in new tab) (2026) analysis, which also tracks Terra at 2x and Luna at 2x cheaper than Fable 5.
Why it matters
GPT-5.6 Sol has been in the directory at a Pending verdict since its June 26 government-gated preview. The independent benchmark data the directory was waiting for has now arrived — and Sol has delivered.
Fable 5's subscription-included access expires tonight at 11:59 PM PT. After that, it's credit-only at $10/$50 — Anthropic's most expensive publicly available model pricing. There is no announced return to subscriptions.
GPT-5.6 Sol launches at $5/$30, subscription-included, and now carries a confirmed coding-agent benchmark lead.
The government-gating drama that surrounded Sol's initial preview obscured a simpler competitive reality: when both models are available, one costs dramatically less and solves more real coding tasks. The policy question was always a sideshow. The benchmark is the main event.
For teams building coding agents, the cost-per-task math compounds fast. A 50-task overnight agent run at $22/task (Fable 5) costs $1,100. The same workload on Sol costs roughly $420. Over a month of daily runs, that gap is $20,000.
What changes for you
If you are choosing between GPT-5.6 Sol and Claude Fable 5 for coding-agent workloads, the directional signal is unambiguous. Sol delivers measurably more solved tasks at one-third the cost. Run your own representative repos to confirm, but do not wait for a tie-breaker that already arrived.
If Fable 5 is in your pipeline, audit today. After 11:59 PM PT, every agent loop burns credits at $10/$50 with no metered subscription fallback. Budget accordingly — there is no automatic cap on spend.
For teams that do not need max reasoning on every task: GPT-5.6 Terra (77 on the Coding Agent Index at ~$0.55/task) and Luna (75 at ~$0.21/task) offer even better value for routine coding work.
FAQ
Is DeepSWE a reliable benchmark? DeepSWE tests agents inside real open-source repositories with program-based verifiers — not multiple-choice questions or function-completion tasks. It measures whether a patch actually works in practice. The benchmark spans 113 tasks, 91 repositories, and 5 languages, making it more representative of real-world engineering than SWE-bench Verified, which relies on unit-test pass/fail.
Does this mean Sol is the better model for everything? No. On AA-Briefcase (knowledge-work tasks), Fable 5 still leads with a Rubric Score of 56% vs Sol's 42%. On the Intelligence Index, the two are one point apart. The data says Sol is the better coding agent for the money — not that Sol is universally superior. For analytical reasoning or presentation-quality outputs, Fable 5 remains competitive.
What about GPT-5.6 Terra and Luna? Both score respectably on the Coding Agent Index — Terra at 77 and Luna at 75 — at even lower per-task cost. If your coding tasks don't require max reasoning effort, the cheaper variants may offer better value still. The entire GPT-5.6 family now defines a new Pareto frontier on intelligence-vs-cost charts.
GPT-5.6 Sol leads the DeepSWE coding-agent benchmark at 72-73% vs Fable 5's 70% at one-third the cost, and tops the Artificial Analysis Coding Agent Index at 80 points. These are the independent benchmarks the directory was waiting for — the verdict moves from Pending to Recommended.
What to do
- 1 If your coding-agent workload is budget-sensitive, benchmark GPT-5.6 Sol against Fable 5 on your own repos — the public data favors Sol on both performance and cost
- 2 Don't let Fable 5's credit-only pricing surprise you — tonight is the last night of subscription-included access
Affected tools & models
Never need to catch up again
The weekly delta — only verdict changes and act-now items. No digest filler.