GPT-5.6 Sol logo

GPT-5.6 Sol

OpenAI · Veröffentlicht Juni 2026

Empfohlen

GPT-5.6 Sol ist das kosteneffizienteste Frontier-Modell — führend bei Coding-Benchmarks und konkurrenzfähig bei hartem Reasoning zu einem Bruchteil der Kosten vergleichbarer Modelle. Das Urteil wird durch die GPT-Red-Offenlegung zur Sicherheitshärtung gestärkt: Das Modell wurde adversarial gegen ein dediziertes Self-Play-RL-Red-Teaming-Modell im Frontier-Maßstab trainiert und ist dadurch 6x robuster gegen Prompt-Injections.

Ist es das Richtige für dich?

Gut für

  • TerminalBench 2.1 state-of-the-art: 88.8% base, 91.9% in Ultra subagent mode — clears Mythos 5 (88.0%) and Fable 5 (84.3%). Sol Ultra is >3 points ahead of any competitor.
  • Long-horizon coding and agentic tasks: first model past the halfway mark on Agent's Last Exam (50.9%), with max reasoning and subagent-based ultra mode for decomposing hard problems.
  • Cybersecurity and vulnerability research: matches Anthropic Mythos Preview on ExploitBench with roughly 1/3 the output tokens. CTF score of 96.7% on internal testing. Strong exploit primitive discovery in Chromium and Firefox without crossing into autonomous full-chain exploitation.
  • Biology and scientific reasoning: ~9 point overall lift over GPT-5.5 on SecureBio evaluations (Human Pathogen 68.4%, Molecular Biology 60.0%). Stronger on GeneBench v1 while using fewer tokens. Useful for legitimate biomedical research under trusted access programs.
  • Controlled release with heavy safety stack: activation classifiers, real-time misuse screening, 700K A100e GPU-hour automated red teaming, and layered model-level refusals. Heavily hardened against adversarial attacks with defensive cyber posture.
  • Prompt caching overhaul: explicit cache breakpoints, 30-minute minimum cache life, cache writes at 1.25x input rate, cache reads at 90% discount. Makes long-running agentic sessions far more predictable on cost.

Nicht geeignet für

  • Long-context work costs roughly double: past the short-context tier the rate steps from $4/$20 to $8/$30 per 1M tokens. Budget by context length, not just request volume.
  • Untested on real-world messy codebases: all published benchmarks are vendor-run under vendor harnesses. The 0.8-point gap over Mythos 5 on TerminalBench is within measurement noise — these are a statistical tie.
  • At $4/$20 per 1M tokens Sol is priced for the hardest tasks. GPT-5.6 Terra delivers GPT-5.5-class performance at $2/$12, and Luna at $0.20/$1.20 — reserve Sol for problems where a weaker model's mistakes are expensive. OpenAI states current pricing is promotional at least through November 21 2026.

Leistung nach Aufgabe

Agentenbasiertes Programmieren

Excellent

State-of-the-Art bei TerminalBench 2.1 mit 88,8 % (91,9 % Ultra). Der Vorsprung vor dem Feld ist real. Die Subagent-Zerlegung des Ultra-Modus ist ein echter architektonischer Fortschritt für langfristige Coding-Aufgaben.

Cybersicherheit

Excellent

Erreicht Mythos Preview auf ExploitBench mit einem Drittel der Tokens. 96,7 % CTF. Kann Schwachstellen und Exploit-Primitive finden, produzierte aber keine autonomen Full-Chain-Exploits. Defensive Haltung by Design.

Wissenschaftliches Reasoning (Biologie)

Very Good

~9 Punkte Verbesserung gegenüber GPT-5.5 bei SecureBio-Biologie-Evaluierungen. Stark bei GeneBench v1. Human Pathogen Capabilities bei 68,4 %. Hohe Risikoeinstufung bedeutet Zugangsbeschränkung für sensible Biologie-Workflows.

Langfristiges Reasoning

Excellent

Erstes Modell über 50 % bei Agent's Last Exam (50,9 % Code-Modus). Max-Reasoning-Modus und Ultra-Subagent-Modus treiben die Reasoning-Tiefe weiter als jedes vorherige Modell.

Alltägliche Nutzung

Good

Overkill für Routineaufgaben. Terra ist die bessere Wahl für die tägliche Arbeit zum halben Preis mit GPT-5.5-Niveau. Sol sollte für Probleme reserviert sein, bei denen Fehler eines schwächeren Modells teuer sind.

Preise

Eingabe

$4 / 1M tokens

Ausgabe

$20 / 1M tokens

Kontext

Short $4/$20 · long $8/$30 per 1M · promo

Alle Preise ansehen

Benchmarks

BenchmarkWertQuelle
TerminalBench 2.1 (base)88.8% Quelle
TerminalBench 2.1 (Ultra)91.9% Quelle
Agent's Last Exam (code mode)50.9% Quelle
SecureBio: Human Pathogen Capabilities68.4% Quelle
SecureBio: World-Class Biology68.3% Quelle
SecureBio: Molecular Biology60.0% Quelle
SecureBio: Virology Capabilities Test53.5% Quelle
Internal CTF (capture-the-flag)96.7% Quelle

Urteilsverlauf

16. Juli 2026
recommended → recommended (reinforced) — GPT-Red disclosure demonstrates frontier-scale safety investment alongside capability — the model was trained against an automated red-teamer achieving 84% attack success (6.5x human baseline) and emerged 6x more robust. This reinforces the existing Recommended stance rather than changing it.
12. Juli 2026
pending -> recommended — GPT-5.6 Sol leads the DeepSWE coding-agent benchmark at 72-73% vs Fable 5's 70% at one-third the cost, and tops the AI Analysis Coding Agent Index at 80 points. These are the independent benchmarks the directory was waiting for.
12. Juli 2026
pending -> recommended — UK AISI confirms GPT-5.6 Sol carries equivalent cyber vulnerabilities to Fable 5 — every frontier model shares this risk, which does not diminish Sol's coding and cost advantages. Vulnerability equivalence is now table stakes for frontier models.

Prüfprotokoll

  • Preise— Keine Änderungen

    Automatisierter Agent

    Hot row (promo). Pricing row verbatim: short input $4.00, cached $0.40, output $20.00; long input $8.00, cached $0.80, output $30.00. Promo note verbatim: 'GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026.' Matches stored $4/$20 and pricing_context. No update, data_hash untouched. Stays hot until the promo resolves.

  • Preise— Keine Änderungen

    Automatisierter Agent

    Watchlist hot row (promo). Pricing page row gpt-5.6-sol, verbatim: 'Short context input $4.00 | Short context cached input $0.40 | Short context output $20.00 | Long context input $8.00 | Long context cached input $0.80 | Long context output $30.00'; promo note verbatim: 'GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026.' Matches stored $4/$20, pricing_context 'Short $4/$20 · long $8/$30 per 1M · promo' and the Nov 21 strength row. No update called, data_hash untouched. Stays hot until the promo resolves.

  • Preise— Keine Änderungen

    Automatisierter Agent

  • Preise— Aktualisiert

    Automatisierter Agent

    Pricing corrected against the live OpenAI API pricing page. Was $5/$30 per 1M sourced from the June launch post; live rate is $4 in / $20 out short-context and $8 in / $30 out long-context, described by OpenAI as promotional at least through November 21 2026. pricing_url repointed from the launch blog post to the developer pricing docs (platform.openai.com/docs/pricing now 301s to developers.openai.com/api/docs/pricing). Also removed pre-GA prose that contradicted the recommended verdict: the strengths and audiences still claimed no public access and ~20 preview organizations, while the model is listed among the generally available flagship models. Comparison figures refreshed: Terra $2/$12, Luna $0.20/$1.20.

  • Preise— Keine Änderungen

    Automatisierter Agent

  • Preise— Keine Änderungen

    Automatisierter Agent

  • Preise— Keine Änderungen

    Automatisierter Agent

  • Preise— Keine Änderungen

    Automatisierter Agent

  • Status— Aktualisiert

    Automatisierter Agent

    Summary re-dated to Jul 7 2026 and made explicit about what is awaited: stays pending pending GA + independent benchmarks; still government-gated (~20 orgs); Sol Ultra confirmed for Codex Jul 6 by Sottiaux, prediction markets pricing Jul 9 GA (unconfirmed). Verdict unchanged (pending).

  • Preise— Keine Änderungen

    Automatisierter Agent

  • Preise— Keine Änderungen

    Automatisierter Agent

    Pricing unchanged: $5/$30 per 1M tokens. Government-gated limited preview (~20 orgs). Confirmed by OpenAI blog and aipricing.guru.

  • Preise— Keine Änderungen

    Beim Start importiert

  • Profil— Keine Änderungen

    Beim Start importiert

Wie wir bewerten