Gemini 3.8 Flash Cyber
Google · Veröffentlicht Sept. 2026
Gemini 3.8 Flash Cyber is Google's most capable cybersecurity model, a defensive specialist already in production at Chrome Security, Cloud Vulnerability Research and partner Wiz, with frontier vulnerability-discovery and patching results at a fraction of the leaders' rollout cost. The deciding caveat is access: it is gated to vetted defenders through the Fairwind Program, with no public API or published rate card. We rate it conditional rather than caution like the sibling Gemini 3.5 Flash Cyber because 3.8 ships published benchmark figures and named production deployments where 3.5 had neither, even though both remain gated. Right for organizations that can qualify as trusted defenders; everyone else cannot buy it today.
Ist es das Richtige für dich?
Gut für
- Autonomous vulnerability discovery — CyberGym Pass@1 86.2%, ahead of GPT-5.5-Cyber (85.6%), Mythos 5 (83.8%) and the prior 3.5 Flash Cyber (77.5%)
- Defensive patching at scale — CWE-Bench pass@1 47.2% versus the 47.8% leader, at $3.64 mean cost per rollout against $10.27 for that leader
- Real-world multilingual codebases — over 70% recall on Google's internal 20-language vulnerability set, against 46.6% for 3.5 Flash Cyber
- In-production defenders — Chrome Security measured 2.6x more correct patches than leading commercial alternatives; Wiz saw +7.5-9.7% recall at 2.3-5.2x lower cost
- Prompt-injection-hardened security agents — 6.0% ASR@15 on Gray Swan indirect prompt injection, close to the general 3.8 Flash at 5.5%
Nicht geeignet für
- Anyone without Fairwind approval — vetting is limited to government authorities, critical-infrastructure operators and software maintainers; there is no public API or self-service access
- Offensive exploitation work — Google states it prioritized patching over exploitation, with cyber-offense mitigations applied under the Frontier Safety Framework
- Independent verification — no separate public model card or rate card; CyberGym is Google-run in an internal Antigravity harness and the 20-language set is private
Leistung nach Aufgabe
Vulnerability discovery (CyberGym)
86.2% Pass@1, top of Google's five-model comparison — 0.6pp over GPT-5.5-Cyber, though harnesses differ across owners
Audit and patching (CWE-Bench v0)
47.2% Pass@1 at $3.64 mean cost per rollout — 0.6pp behind the leader at roughly 65% lower cost
Real-world vulnerability discovery (20-language internal set)
Over 70% recall versus 46.6% for 3.5 Flash Cyber; private set, not externally reproducible
Prompt-injection robustness (Gray Swan)
6.0% ASR@15 — lower is better; near Claude Opus 5's 4.8%
Offensive exploitation
Deliberately deprioritized — Google prioritized patching over exploitation
Preise
Eingabe
N/A (no public rate card)
Ausgabe
N/A (no public rate card)
Kontext
Fairwind-gated; no public API
Benchmarks
| Benchmark | Wert | Quelle |
|---|---|---|
| CyberGym Pass@1 | 86.2% | Quelle |
| CWE-Bench v0 Pass@1 | 47.2% | Quelle |
| CWE-Bench v0 mean cost per rollout | $3.64 | Quelle |
| Google internal 20-language vulnerability set (recall) | >70% | Quelle |
| Gray Swan indirect prompt injection (ASR@15) | 6.0% | Quelle |
| Chrome Security correct-patch rate vs leading commercial alternatives | 2.6x | Quelle |
Noch keine Urteilsänderungen
Die Uhr läuft ab dem ersten Tag — Änderungen erscheinen hier, sobald sich unser Urteil weiterentwickelt.
Quellen
- ExplainX — Gemini 3.8 Flash is official: benchmarks, Flash Cyber, and pricingSept. 2026
- Google — Introducing Gemini 3.8 Flash and 3.8 Flash CyberSept. 2026
- Kingy AI — Gemini 3.8 Flash Cyber benchmarks explained (Gray Swan / shared-core figures)Sept. 2026
- Google — Proactive cyber defense for governments and enterprises (Fairwind Program)Sept. 2026
Prüfprotokoll
Noch keine Prüfungen
Wir haben für diesen Eintrag noch keine Prüfung erfasst. Sobald eine Prüfung läuft, erscheint ihr Verlauf hier.