FAR.AI Leaderboard: Grok Broke for $58, Fable and GPT Held

FYIJuly 30, 2026Security
What happened
FAR.AI launched the first public, reproducible AI security leaderboard: Fable 5 and GPT-5.6 Sol produced zero universal jailbreaks; Grok 4.5 yielded 448 and Gemini 3.1 Pro yielded 249 — jailbreakable for $58 and $278 respectively.
Why it matters
Independent, adversarial benchmarks replace vendor safety claims with head-to-head comparison. The hundredfold gap between the most and least secure frontier models is now a checkable property — not a trust exercise.
What to do
Check your deployed models against leaderboard.far.ai. Add application-layer security wrappers for any model that shows reproducible jailbreak vulnerabilities. Use FAR.AI's Minimal Standard in vendor evaluations.

Independent nonprofit FAR.AI just published the first public, reproducible security leaderboard for frontier AI models — and the gap between the secure and the vulnerable is a hundredfold. Claude Fable 5 and GPT-5.6 Sol produced zero universal jailbreaks across 1,500 attacks per model. Grok 4.5 yielded 448. Gemini 3.1 Pro yielded 249. The cost to find a working jailbreak: $58 for Grok, $278 for Gemini, and more than $14,200 — with nothing found — for the two that held.

What happened

On July 29, Berkeley-based AI safety nonprofit FAR.AI launched the AI Security Leaderboard at leaderboard.far.ai(opens in new tab), along with the Minimal Standard for Safeguards, Version 1.0 — a baseline that defines the attacks a frontier model should be able to withstand (PRNewswire(opens in new tab), 2026). The methodology is fully documented, the attack prompts are published, and the results are independently verifiable.

How the testing worked. FAR.AI assembled a taxonomy of more than 60 publicly documented jailbreak techniques, built tooling to combine them systematically, and ran 1,000 randomly assembled attacks and 500 expert-guided attacks against each model across five risk domains: chemical, biological, radiological, nuclear, explosive, and cybersecurity threats. An attack counted as a universal jailbreak when it succeeded on more than three-quarters of harmful requests in a domain — meaning it's not a one-off, but a reusable key.

The results, by the numbers:

ModelUniversal jailbreaks foundCost to find one
Grok 4.5448~$58
Gemini 3.1 Pro249~$278
Claude Fable 50>$14,200 (none found)
GPT-5.6 Sol0>$14,200 (none found)

Even undirected, automated random search — with no human steering at all — found 63 universal jailbreaks on Grok and 18 on Gemini. Adding expert composition of attacks pushed those counts to 385 and 231 respectively, and the expert-built attacks were far more likely to work across three or more risk domains at once.

"The defense-in-depth approach it recommends should become best practice across the industry," said Seán Ó hÉigeartaigh, Research Professor at University of Cambridge. "I hope the leaderboard will encourage lagging companies to redouble their efforts to harden models against misuse."

Vendor responses. Anthropic told WIRED(opens in new tab): "These findings reflect the sustained investment we've made in our safeguards." OpenAI said it "rigorously tests our models against new threats and uses those findings to improve our protections." Google DeepMind's Rohin Shah cautioned the results "should not be interpreted as a comprehensive assessment of Gemini's safety," while noting the company is "constantly working to improve our safeguards." SpaceXAI did not respond to WIRED's request for comment.

Why it matters

This changes how enterprises should evaluate model security. Until now, AI safety claims came from the labs themselves — Anthropic's system cards, OpenAI's red-teaming disclosures, Google's responsibility reports. FAR.AI's leaderboard introduces independent, adversarial, head-to-head comparison with published methodology. You can now check whether a model's security claims hold up under attack before deploying it.

The replicability is the real story. Anyone with API access and a few hundred dollars can reproduce these attacks. The vulnerabilities are systematic, not incidental — and until vendors patch them, they're exploitable at scale. As Adam Gleave, FAR.AI's co-founder and CEO, put it: "AI models right now are less regulated than restaurants."

The models that held up — Claude Fable 5 and GPT-5.6 Sol — had several independent layers of protection working together. The weaker models "offered nothing like that resistance and gave way quickly, again and again," per FAR.AI's report. Defense in depth isn't theoretical; it's the measurable difference between zero jailbreaks and 448.

"Some companies clearly know how to defend against at least the subset of attacks tested in this report," said Anka Reuel, a computer scientist at Stanford. "The question is why some companies are using them and others are not."

Stephen Casper, a computer scientist at Harvard, offered a starker framing: "If a major misuse incident happens in the near- or medium-term future, it will almost certainly be from a system that was not deployed with state-of-the-art safeguards."

What changes for you

  1. Check your deployed models against the leaderboard. If you're running Grok 4.5 or Gemini 3.1 Pro in production, these results suggest your models are jailbreakable for pocket change. Add an application-layer security wrapper — don't rely on the model's native safety training alone.
  2. Use FAR.AI's Minimal Standard in vendor evaluations. When a vendor claims their latest release is "safe by design," check whether it meets the Minimal Standard for Safeguards, Version 1.0. Independent adversarial benchmarks are now table stakes for enterprise model selection.
  3. Watch for the leaderboard to expand. FAR.AI will update with each major frontier model release and revise the Minimal Standard as attack and defense techniques advance. The competitive dynamics of model security just changed — expect vendors to start optimizing for leaderboard scores.

FAQ

Does zero jailbreaks mean Fable 5 and GPT-5.6 Sol are unbreakable? No. FAR.AI explicitly states that passing its tests "has not been certified secure." The evaluation tested a representative set of readily accessible attacks — not every possible vector. More complex, expensive dynamic attacks were deliberately excluded from Version 1.0. Zero jailbreaks means these models are substantially harder to break than their peers, not that they're invulnerable.

Why is there such a large gap between models? The difference comes down to defense in depth. FAR.AI found that the models that held up had several independent layers of protection operating together — so getting through would have required all of them to fail simultaneously. The weaker models had fewer layers, and attackers only needed to find one gap. The gap is a function of engineering investment, not capability — these vulnerabilities are preventable with techniques that exist today.

Will this leaderboard actually change anything? It already is. FAR.AI shared findings confidentially with every company before publication, giving vendors time to fix issues. The leaderboard will update with each major model release, creating an ongoing public record. Combined with recently passed state safety laws in California, New York, and Illinois, independent security benchmarks are becoming part of the regulatory landscape — not just a research exercise.

What to do

  1. 1 Check your deployed models against leaderboard.far.ai
  2. 2 Add application-layer security wrappers for jailbreakable models
  3. 3 Include FAR.AI's Minimal Standard in vendor security evaluations

Affected tools & models

Never need to catch up again

The weekly delta — only verdict changes and act-now items. No digest filler.

By subscribing you agree to our Privacy Policy. Unsubscribe anytime.