GPT-5.5-Cyber Finds Hundreds of OSS Bugs in First Week

OpenAI logoOpenAIImportantJune 23, 2026Models
What happened
OpenAI's GPT-5.5-Cyber found hundreds of real vulnerabilities — including 8 Linux kernel PoCs, 34 FreeBSD CVEs, and 5 Chrome V8 bugs — in a five-day sprint across 19 open-source projects, with Trail of Bits engineers validating every finding.
Why it matters
This is the first frontier model deliberately deployed as a defense weapon at internet scale. It's not just coding — it's finding exploitable bugs faster than any human team, with expert triage that actually reduces the burden on maintainers.
What to do
Open-source maintainers should apply to participate via Trail of Bits. Security teams should evaluate GPT-5.5-Cyber against internal tooling. If this proves sustained, revisit the GPT-4o verdict.

GPT-5.5-Cyber, deployed as part of OpenAI's Patch the Planet initiative with Trail of Bits, found hundreds of real vulnerabilities in a single five-day sprint — 8 Linux kernel PoCs, 34 FreeBSD CVEs, 5 Chrome V8 bugs, and a WebAssembly flaw that caused five teams to withdraw from Pwn2Own Berlin before the competition started. This is the first time a frontier AI model has been deliberately weaponized for open-source defense at internet scale. It worked.

What happened

On June 22, 2026, OpenAI expanded its Daybreak cybersecurity initiative with three releases: the full version of GPT-5.5-Cyber, restricted to trusted defenders via the Trusted Access for Cyber program; a Codex Security plugin for vulnerability scanning inside developer workflows; and Patch the Planet — an initiative with Trail of Bits to find, validate, and patch vulnerabilities in the open-source software the world depends on.

Trail of Bits committed 25 engineers — roughly a fifth of its workforce — to work full-time with GPT-5.5-Cyber across 19 open-source projects. Every AI-generated finding was manually reviewed before reaching maintainers. The initial sprint produced hundreds of discovered bugs, 64 pull requests, and 51 issues filed, with many more still under coordinated disclosure (OpenAI, 2026(opens in new tab)).

The numbers from week one

ProjectFindingCount
Linux KernelKernel pointer info-leak PoCs8
Linux KernelLocal privilege escalation exploits24
FreeBSDConfirmed vulnerabilities34
FreeBSDLocal privilege escalation PoCs7
Chrome V8Exploitable bugs (3 fixed within days)5
Safari/WebKitExploitable vulnerabilities10+
FirefoxWebAssembly CVE-2026-8390, patched before Pwn2Own1
OpenBSD23-year-old use-after-free in SysV semaphore1
dnsmasqIndependently found 4 of 6 CVEs in v2.92rel24
HTTP/2 BombDoS affecting NGINX, Apache, IIS, Pingora880K+ sites

On CyberGym, GPT-5.5-Cyber scored 85.6%, beating GPT-5.5 (81.8%) and Anthropic's Mythos 5 (83.8%) (MLQ, 2026(opens in new tab)).

The Firefox finding was the most dramatic: a WebAssembly vulnerability (CVE-2026-8390) was patched two days before Pwn2Own Berlin. Five of six registered Firefox entries withdrew. No Firefox exploit was demonstrated at the competition (WIRED, 2026(opens in new tab)).

Why it matters

Open-source maintainers are drowning. AI-generated vulnerability reports flood their queues with false positives, and the same limited volunteers are asked to triage more noise than ever. Patch the Planet solves the actual bottleneck: Trail of Bits engineers reproduce the evidence, filter false positives, reassess severity, and submit vetted patches — reducing the burden on maintainers, not adding to it.

As Trail of Bits CEO Dan Guido put it: "Patch the Planet is an internet-scale effort to help open-source software get ahead of AI bug-hunting tools. But it's also an effort to help the open-source community see the benefits and not just the downsides of AI coding tools" (Trail of Bits, 2026(opens in new tab)).

The Five Eyes intelligence alliance took the extraordinary step of issuing a joint statement the same day: "Frontier AI models are anticipated to exceed current industry expectations, fundamentally transforming both offensive and defensive cyber capabilities. The timeline is not years, it is months."

The capability signal is unmistakable. A frontier model, deliberately optimized for cyber offense-turned-defense, found real, exploitable vulnerabilities across every layer of the software stack — from browsers to kernels to network infrastructure — in days. GPT-5.5-Cyber is not publicly available; it requires Trusted Access for Cyber approval, which mandates phishing-resistant authentication and blocks credential theft, stealth, persistence, and malware deployment. The model is explicitly restricted to defensive workflows.

The standout engineering detail: Trail of Bits engineers used GPT-5.5-Cyber to build a complete fuzzing lab in less than a day — work that would normally take "at least several weeks." The model made autonomous decisions about coverage expansion, build variants, and candidate filtering, with engineers setting objectives and refining prompts.

Initial participants — cURL, NATS Server, pyca/cryptography, Sigstore, aiohttp, the Go project, freenginx, Python, and python.org — receive six months of ChatGPT Pro, conditional Codex Security access, and API credits for core development. More than 30 projects have committed.

If this becomes a sustained capability rather than a launch-day demonstration, our GPT-4o and broader GPT-5.x verdicts deserve a hard re-examination. The defense community just got a new tool, and the implications for vulnerability discovery timelines are profound.

What changes for you

If you maintain open-source software: Apply to participate. More than 30 projects have already committed and receive ChatGPT Pro, Codex Security access, and API credits. The barrier is vetting — Trail of Bits engineers handle the triage, so you won't face a flood of unvalidated AI reports.

If you're on a security team: GPT-5.5-Cyber via the Trusted Access for Cyber program is the new bar for automated vulnerability discovery. Start evaluating it against your internal tooling. The Five Eyes statement means regulatory requirements around AI cyber capabilities are accelerating — compliance teams should track this.

If you're evaluating AI tools: Watch whether this becomes a sustained capability. A one-week sprint is impressive; a persistent program that keeps finding bugs month over month would change the value proposition of frontier models entirely.

FAQ

Can I access GPT-5.5-Cyber? Not without approval. GPT-5.5-Cyber is restricted to trusted defenders via the Trusted Access for Cyber program, which requires phishing-resistant authentication and explicitly blocks offensive capabilities. If you maintain open-source infrastructure, apply through Patch the Planet for mediated access via Trail of Bits.

How does this compare to Anthropic's Mythos 5? On CyberGym, GPT-5.5-Cyber scores 85.6% vs Mythos 5's 83.8%. But the real difference is deployment: GPT-5.5-Cyber is actively finding bugs in production infrastructure with expert human validation. Mythos 5's real-world cyber performance hasn't been demonstrated at comparable scale yet.

Is this a one-off launch demo? The pressure test is sustainability. A five-day sprint with 25 Trail of Bits engineers produced extraordinary results, but the question is whether the pipeline can maintain velocity across 30+ committed projects. OpenAI is betting that it can — and the initial infrastructure (fuzzing harnesses, CVE analysis pipelines, differential testing systems) is built for reuse.

What to do

  1. 1 Open-source maintainers: apply to Patch the Planet via Trail of Bits for free security review with expert-validated findings
  2. 2 Security teams: evaluate GPT-5.5-Cyber against internal vulnerability discovery tooling
  3. 3 Track whether this becomes a sustained program — a persistent capability would change frontier model value assessment

Affected tools & models

GPT-5.5-CyberGPT-4o

Never need to catch up again

The weekly delta — only verdict changes and act-now items. No digest filler.

By subscribing you agree to our Privacy Policy. Unsubscribe anytime.