UK AISI: GPT-5.6 Sol, Fable 5 Share Identical Cyber Risk

UK AISI logoUK AISIVerdict changedJuly 12, 2026Policy & Regulation
What happened
The UK's AI Security Institute confirmed GPT-5.6 Sol has the same cyber vulnerabilities as Fable 5 — and found a universal jailbreak potentially more serious than the one that triggered Fable 5's 19-day ban.
Why it matters
The US banned Fable 5 for 19 days over the same risk profile. GPT-5.6 Sol launched globally with no restrictions. The government-gating asymmetry now has an official safety evaluation as evidence — published in OpenAI's own System Card.
What to do
Re-evaluate any procurement or policy decisions that assume one frontier model is inherently safer than another. Test models directly — don't rely on government gating as a safety proxy.

The UK's AI Security Institute confirmed on July 9 — in findings published in OpenAI's own GPT-5.6 System Card(opens in new tab) — that GPT-5.6 Sol carries the same cyber vulnerabilities as Anthropic's Fable 5. The jailbreak AISI discovered may actually be worse: the researchers described it as "universal" and said it enables autonomous exploits, not just vulnerability identification. The finding exposes a regulatory asymmetry that can no longer be justified by risk: Fable 5 was banned from foreign access for 19 days over the same vulnerability profile, while GPT-5.6 Sol launched globally with no restrictions at all.

What happened — AISI's GPT-5.6 Sol jailbreak findings

British government researchers at the UK's AI Security Institute (AISI) identified what they called "universal jailbreaks in the cyber domain" in GPT-5.6 Sol, including jailbreaks that "allowed for long-form agentic task completion in domains like vulnerability discovery and exploit development" (GPT-5.6 System Card(opens in new tab), July 9, 2026). In plain terms: they jailbroke the model and made it autonomously find software vulnerabilities and hack systems.

The jailbreak is described as "general-purpose" — it enables standalone exploits, not just vulnerability identification. The AISI noted this is potentially more serious than the jailbreak found in Fable 5, which the Amazon discovery team described as unlocking only vulnerability finding, not autonomous exploitation (Fortune(opens in new tab), July 10, 2026). That Fable 5 jailbreak triggered a 19-day US government ban in June.

OpenAI acknowledged the findings but pushed back on their severity for ordinary users. The company noted AISI researchers had "privileged access to the internal mechanisms of the system, which allowed them to speed up the hacking process." However, Xander Davies, who leads AISI's red team, posted on X that he believed the jailbreaks his team discovered "are still findable without this access, just slower" (Fortune(opens in new tab), 2026).

OpenAI referenced its GPT-5.6 launch blog, which states that "perfect security does not exist" and that the company takes a "multi-layered" approach including continuous monitoring and rapid remediation. Margaret Cunningham, VP of security at DarkTrace and NIST specialist collaborator, offered the most grounded take: "My concern is less that one model was jailbroken and more that offensive discovery is speeding up while defense still depends on very human processes" (Fortune(opens in new tab), 2026).

Stanislav Fort, chief scientist at AI cybersecurity startup AISLE and a former researcher at both Anthropic and Google DeepMind, summarized the uncomfortable reality: "Every deployed model right now almost certainly has undiscovered jailbreaks, so this is sadly true of everything, not just GPT-5.6" (Fortune(opens in new tab), 2026).

Why it matters

The containment-through-restriction argument — the idea that we can make the world safer by locking down specific models — has now been empirically invalidated by the very safety institute designed to validate it.

The numbers tell the story:

ModelVulnerability profileGovernment actionCurrent access
Fable 5Cyber exploit capabilityBanned 19 days (Jun 12–Jul 1)Credit-only, $10/$50 per 1M tokens
GPT-5.6 SolSame vulnerabilities (AISI confirmed)No restrictionsSubscription-included, $5/$30 per 1M tokens

Both models pose equivalent cyber risk. Only one was shut down. The other launched globally with subscription-included access at a lower price point.

This asymmetry matters because it exposes the government's export-control framework as reactive and inconsistent rather than risk-calibrated. When the US government claims to restrict models based on national security risk, the AISI data — published in OpenAI's own System Card — now shows the same risk profile receiving opposite treatment. Either the standard is vulnerability equivalence — in which case both should have been restricted — or the Fable 5 ban was a mistake. The government has done neither.

Lennart Heim, an AI policy researcher, captured the absurdity with a post on X: "good thing amazon didn't report this one to the white house" — a reference to how the Trump administration learned of the Fable 5 jailbreak from Amazon's discovery team. One former AI policy advisor told Fortune: "What we are seeing recently creates uncertainty that is damaging in the least and potentially raises the question of whether, intentional or not, the U.S. is applying an inconsistent standard to different AI labs" (Fortune(opens in new tab), 2026).

For organizations that rely on government gating as a safety signal, this finding should trigger an immediate reassessment. If two models with identical vulnerability profiles receive different regulatory treatment, the regulatory treatment itself is not a reliable indicator of safety.

What changes for you

Re-evaluate procurement decisions. Don't assume a model that cleared government review is inherently safer than one that didn't. The AISI data shows they're equivalently vulnerable — only the regulatory treatment differs.

Test directly. If your organization uses AI models for security-sensitive work, test both models against your own threat model rather than relying on government approval as a proxy for safety. The Stanislav Fort observation — that every deployed model almost certainly has undiscovered jailbreaks — should be your baseline assumption.

Watch the policy response. Two equally vulnerable models receiving opposite regulatory treatment creates a policy tension that can't hold. Either the restrictions tighten on all frontier models, or the framework collapses under its own inconsistency. Both scenarios change the procurement calculus.

FAQ

Does this mean GPT-5.6 Sol is unsafe? Not uniquely — that's the point. The AISI finding confirms that vulnerability equivalence is table stakes for frontier models. GPT-5.6 Sol is no more and no less dangerous than Fable 5 on cybersecurity grounds. Every model at this tier carries undetected jailbreaks.

What should my security team do? Treat all frontier models as equivalently risky for cybersecurity-sensitive workloads. Build your own testing into procurement workflows rather than outsourcing safety judgments to government gating decisions.

pendingprevious pick
recommendednew pick

UK AISI confirms GPT-5.6 Sol carries equivalent cyber vulnerabilities to Fable 5 — every frontier model shares this risk, which does not diminish Sol's coding and cost advantages. Vulnerability equivalence is now table stakes for frontier models; the verdict moves from Pending to Recommended.

What to do

  1. 1 Re-evaluate any procurement or policy decisions that assume one frontier model is inherently safer than another
  2. 2 If your organization uses AI models for security-sensitive work, test both — don't rely on government gating as a safety proxy

Affected tools & models

Never need to catch up again

The weekly delta — only verdict changes and act-now items. No digest filler.

By subscribing you agree to our Privacy Policy. Unsubscribe anytime.