GPT-5.6 Sol logo

GPT-5.6 Sol

OpenAI · Lançado 06/2026

Recomendado

O GPT-5.6 Sol é o modelo de fronteira mais custo-efetivo, liderando em benchmarks de codificação e competitivo em raciocínio complexo a uma fração do custo dos concorrentes. O veredicto é reforçado pela divulgação do endurecimento de segurança GPT-Red — o modelo foi treinado adversarialmente contra um modelo dedicado de red-teaming por RL de auto-jogo em escala de fronteira, tornando-o 6x mais robusto a injeções de prompt.

É adequada para ti?

Bom para

  • TerminalBench 2.1 state-of-the-art: 88.8% base, 91.9% in Ultra subagent mode — clears Mythos 5 (88.0%) and Fable 5 (84.3%). Sol Ultra is >3 points ahead of any competitor.
  • Long-horizon coding and agentic tasks: first model past the halfway mark on Agent's Last Exam (50.9%), with max reasoning and subagent-based ultra mode for decomposing hard problems.
  • Cybersecurity and vulnerability research: matches Anthropic Mythos Preview on ExploitBench with roughly 1/3 the output tokens. CTF score of 96.7% on internal testing. Strong exploit primitive discovery in Chromium and Firefox without crossing into autonomous full-chain exploitation.
  • Biology and scientific reasoning: ~9 point overall lift over GPT-5.5 on SecureBio evaluations (Human Pathogen 68.4%, Molecular Biology 60.0%). Stronger on GeneBench v1 while using fewer tokens. Useful for legitimate biomedical research under trusted access programs.
  • Controlled release with heavy safety stack: activation classifiers, real-time misuse screening, 700K A100e GPU-hour automated red teaming, and layered model-level refusals. Heavily hardened against adversarial attacks with defensive cyber posture.
  • Prompt caching overhaul: explicit cache breakpoints, 30-minute minimum cache life, cache writes at 1.25x input rate, cache reads at 90% discount. Makes long-running agentic sessions far more predictable on cost.

Não recomendado para

  • Long-context work costs roughly double: past the short-context tier the rate steps from $4/$20 to $8/$30 per 1M tokens. Budget by context length, not just request volume.
  • Untested on real-world messy codebases: all published benchmarks are vendor-run under vendor harnesses. The 0.8-point gap over Mythos 5 on TerminalBench is within measurement noise — these are a statistical tie.
  • At $4/$20 per 1M tokens Sol is priced for the hardest tasks. GPT-5.6 Terra delivers GPT-5.5-class performance at $2/$12, and Luna at $0.20/$1.20 — reserve Sol for problems where a weaker model's mistakes are expensive. OpenAI states current pricing is promotional at least through November 21 2026.

Desempenho por tarefa

Codificação agêntica

Excellent

Estado da arte no TerminalBench 2.1 com 88,8% (91,9% Ultra). A diferença sobre o campo é real. A decomposição por subagentes do modo Ultra é um avanço arquitetónico genuíno para tarefas de codificação de longo horizonte.

Cibersegurança

Excellent

Iguala o Mythos Preview no ExploitBench com 1/3 dos tokens. 96,7% CTF. Pode encontrar vulnerabilidades e primitivas de exploit mas não produziu autonomamente exploits de cadeia completa. Postura defensiva por design.

Raciocínio científico (biologia)

Very Good

Aumento de ~9 pontos sobre o GPT-5.5 nas avaliações de biologia SecureBio. Forte no GeneBench v1. Human Pathogen Capabilities a 68,4%. Classificação de alto risco significa acesso restrito para fluxos de biologia sensíveis.

Raciocínio de longo horizonte

Excellent

Primeiro modelo a ultrapassar 50% no Agent's Last Exam (50,9% modo código). Modo de raciocínio máximo e modo ultra com subagente empurram a profundidade de raciocínio mais longe que qualquer modelo anterior.

Uso geral diário

Good

Exagerado para tarefas de rotina. Terra é a melhor escolha para trabalho diário a metade do custo com capacidade de classe GPT-5.5. Sol deve ser reservado para problemas onde os erros de um modelo mais fraco são caros.

Preços

Entrada

$4 / 1M tokens

Saída

$20 / 1M tokens

Contexto

Short $4/$20 · long $8/$30 per 1M · promo

Ver preços completos

Testes de desempenho

TestePontuaçãoFonte
TerminalBench 2.1 (base)88.8% Fonte
TerminalBench 2.1 (Ultra)91.9% Fonte
Agent's Last Exam (code mode)50.9% Fonte
SecureBio: Human Pathogen Capabilities68.4% Fonte
SecureBio: World-Class Biology68.3% Fonte
SecureBio: Molecular Biology60.0% Fonte
SecureBio: Virology Capabilities Test53.5% Fonte
Internal CTF (capture-the-flag)96.7% Fonte

Histórico de veredictos

16/07/2026
recommended → recommended (reinforced) — GPT-Red disclosure demonstrates frontier-scale safety investment alongside capability — the model was trained against an automated red-teamer achieving 84% attack success (6.5x human baseline) and emerged 6x more robust. This reinforces the existing Recommended stance rather than changing it.
12/07/2026
pending -> recommended — GPT-5.6 Sol leads the DeepSWE coding-agent benchmark at 72-73% vs Fable 5's 70% at one-third the cost, and tops the AI Analysis Coding Agent Index at 80 points. These are the independent benchmarks the directory was waiting for.
12/07/2026
pending -> recommended — UK AISI confirms GPT-5.6 Sol carries equivalent cyber vulnerabilities to Fable 5 — every frontier model shares this risk, which does not diminish Sol's coding and cost advantages. Vulnerability equivalence is now table stakes for frontier models.

Registo de verificação

  • Preços— Sem alterações

    Agente automatizado

    Hot row (promo). Pricing row verbatim: short input $4.00, cached $0.40, output $20.00; long input $8.00, cached $0.80, output $30.00. Promo note verbatim: 'GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026.' Matches stored $4/$20 and pricing_context. No update, data_hash untouched. Stays hot until the promo resolves.

  • Preços— Sem alterações

    Agente automatizado

    Watchlist hot row (promo). Pricing page row gpt-5.6-sol, verbatim: 'Short context input $4.00 | Short context cached input $0.40 | Short context output $20.00 | Long context input $8.00 | Long context cached input $0.80 | Long context output $30.00'; promo note verbatim: 'GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026.' Matches stored $4/$20, pricing_context 'Short $4/$20 · long $8/$30 per 1M · promo' and the Nov 21 strength row. No update called, data_hash untouched. Stays hot until the promo resolves.

  • Preços— Sem alterações

    Agente automatizado

  • Preços— Atualizado

    Agente automatizado

    Pricing corrected against the live OpenAI API pricing page. Was $5/$30 per 1M sourced from the June launch post; live rate is $4 in / $20 out short-context and $8 in / $30 out long-context, described by OpenAI as promotional at least through November 21 2026. pricing_url repointed from the launch blog post to the developer pricing docs (platform.openai.com/docs/pricing now 301s to developers.openai.com/api/docs/pricing). Also removed pre-GA prose that contradicted the recommended verdict: the strengths and audiences still claimed no public access and ~20 preview organizations, while the model is listed among the generally available flagship models. Comparison figures refreshed: Terra $2/$12, Luna $0.20/$1.20.

  • Preços— Sem alterações

    Agente automatizado

  • Preços— Sem alterações

    Agente automatizado

  • Preços— Sem alterações

    Agente automatizado

  • Preços— Sem alterações

    Agente automatizado

  • Estado— Atualizado

    Agente automatizado

    Summary re-dated to Jul 7 2026 and made explicit about what is awaited: stays pending pending GA + independent benchmarks; still government-gated (~20 orgs); Sol Ultra confirmed for Codex Jul 6 by Sottiaux, prediction markets pricing Jul 9 GA (unconfirmed). Verdict unchanged (pending).

  • Preços— Sem alterações

    Agente automatizado

  • Preços— Sem alterações

    Agente automatizado

    Pricing unchanged: $5/$30 per 1M tokens. Government-gated limited preview (~20 orgs). Confirmed by OpenAI blog and aipricing.guru.

  • Preços— Sem alterações

    Importado no lançamento

  • Perfil— Sem alterações

    Importado no lançamento

Como avaliamos