GPT-5.6 Sol logo

GPT-5.6 Sol

OpenAI · Released Jun 2026

Recommended

GPT-5.6 Sol is the most cost-effective frontier model, leading on coding benchmarks and competitive on hard reasoning at a fraction of the cost of peers. Verdict reinforced by GPT-Red safety-hardening disclosure — the model was adversarially trained against a dedicated self-play RL red-teaming model at frontier scale, making it 6x more robust to prompt injections.

Is it right for you?

Good for

  • TerminalBench 2.1 state-of-the-art: 88.8% base, 91.9% in Ultra subagent mode — clears Mythos 5 (88.0%) and Fable 5 (84.3%). Sol Ultra is >3 points ahead of any competitor.
  • Long-horizon coding and agentic tasks: first model past the halfway mark on Agent's Last Exam (50.9%), with max reasoning and subagent-based ultra mode for decomposing hard problems.
  • Cybersecurity and vulnerability research: matches Anthropic Mythos Preview on ExploitBench with roughly 1/3 the output tokens. CTF score of 96.7% on internal testing. Strong exploit primitive discovery in Chromium and Firefox without crossing into autonomous full-chain exploitation.
  • Biology and scientific reasoning: ~9 point overall lift over GPT-5.5 on SecureBio evaluations (Human Pathogen 68.4%, Molecular Biology 60.0%). Stronger on GeneBench v1 while using fewer tokens. Useful for legitimate biomedical research under trusted access programs.
  • Controlled release with heavy safety stack: activation classifiers, real-time misuse screening, 700K A100e GPU-hour automated red teaming, and layered model-level refusals. Heavily hardened against adversarial attacks with defensive cyber posture.
  • Prompt caching overhaul: explicit cache breakpoints, 30-minute minimum cache life, cache writes at 1.25x input rate, cache reads at 90% discount. Makes long-running agentic sessions far more predictable on cost.

Not good for

  • Government-gated limited preview with no public access: only ~20 organizations receive access, and general availability is promised 'in the coming weeks' but not guaranteed. You cannot test this model against your own workloads today.
  • Untested on real-world messy codebases: all published benchmarks are vendor-run under vendor harnesses. The 0.8-point gap over Mythos 5 on TerminalBench is within measurement noise — these are a statistical tie.
  • At $5/$30 per 1M tokens, Sol matches GPT-5.5 pricing for a major capability bump — but GPT-5.6 Terra delivers GPT-5.5-class performance at half the cost ($2.50/$15). Sol should be reserved for the hardest tasks only.

How it performs by task

Agentic coding

Excellent

State-of-the-art on TerminalBench 2.1 at 88.8% (91.9% Ultra). The gap over the field is real. Ultra mode's subagent decomposition is a genuine architectural advance for long-horizon coding tasks.

Cybersecurity

Excellent

Matches Mythos Preview on ExploitBench with 1/3 tokens. 96.7% CTF. Can find vulnerabilities and exploit primitives but did not autonomously produce full-chain exploits. Defensive posture by design.

Scientific reasoning (biology)

Very Good

~9-point lift over GPT-5.5 on SecureBio biology evals. Strong on GeneBench v1. Human Pathogen Capabilities at 68.4%. High risk classification means access is gated for sensitive biology workflows.

Long-horizon reasoning

Excellent

First model past 50% on Agent's Last Exam (50.9% code mode). Max reasoning mode and ultra subagent mode push reasoning depth further than any prior model.

Everyday general use

Good

Overkill for routine tasks. Terra is the better fit for daily work at half the cost with GPT-5.5-class capability. Sol should be reserved for problems where a weaker model's mistakes are expensive.

Pricing

Input

$5 / 1M tokens

Output

$30 / 1M tokens

Context

Same rate card as GPT-5.5. Ultra mode costs more.

View full pricing

Benchmarks

BenchmarkScoreSource
TerminalBench 2.1 (base)88.8% Source
TerminalBench 2.1 (Ultra)91.9% Source
Agent's Last Exam (code mode)50.9% Source
SecureBio: Human Pathogen Capabilities68.4% Source
SecureBio: World-Class Biology68.3% Source
SecureBio: Molecular Biology60.0% Source
SecureBio: Virology Capabilities Test53.5% Source
Internal CTF (capture-the-flag)96.7% Source

Verdict history

Jul 16, 2026
recommended → recommended (reinforced) — GPT-Red disclosure demonstrates frontier-scale safety investment alongside capability — the model was trained against an automated red-teamer achieving 84% attack success (6.5x human baseline) and emerged 6x more robust. This reinforces the existing Recommended stance rather than changing it.
Jul 12, 2026
pending -> recommended — GPT-5.6 Sol leads the DeepSWE coding-agent benchmark at 72-73% vs Fable 5's 70% at one-third the cost, and tops the AI Analysis Coding Agent Index at 80 points. These are the independent benchmarks the directory was waiting for.
Jul 12, 2026
pending -> recommended — UK AISI confirms GPT-5.6 Sol carries equivalent cyber vulnerabilities to Fable 5 — every frontier model shares this risk, which does not diminish Sol's coding and cost advantages. Vulnerability equivalence is now table stakes for frontier models.

Verification log

  • Pricing— No changes

    Automated agent

  • Pricing— No changes

    Automated agent

  • Status— Updated

    Automated agent

    Summary re-dated to Jul 7 2026 and made explicit about what is awaited: stays pending pending GA + independent benchmarks; still government-gated (~20 orgs); Sol Ultra confirmed for Codex Jul 6 by Sottiaux, prediction markets pricing Jul 9 GA (unconfirmed). Verdict unchanged (pending).

  • Pricing— No changes

    Automated agent

  • Pricing— No changes

    Automated agent

    Pricing unchanged: $5/$30 per 1M tokens. Government-gated limited preview (~20 orgs). Confirmed by OpenAI blog and aipricing.guru.

  • Profile— No changes

    Imported at launch

  • Pricing— No changes

    Imported at launch

How we evaluate