La semaine où l'IA de défense est devenue opérationnelle

AI Changelog · Numéro 14Week 26 · Jun 22–28, 2026

Le GPT-5.5-Cyber d'OpenAI a trouvé 34 CVE FreeBSD, 8 PoCs du noyau Linux et 5 bugs Chrome V8 en 5 jours — le premier modèle de pointe délibérément déployé comme arme de défense à l'échelle d'Internet. L'alliance Five Eyes a averti le même jour que les cyberattaques par IA sont à des mois, pas des années.

61 announcements scanned · 15 mattered · 0 verdicts changed

À traiter avant le prochain numéro

Re-examine our GPT-4o verdict — Patch the Planet is the reason

GPT-5.5-Cyber found 34 FreeBSD CVEs, 8 Linux kernel PoCs, and 5 Chrome V8 bugs in 5 days, validated by 25 Trail of Bits engineers. If this pipeline sustains month over month, our GPT-4o recommended call needs a hard re-examination. We commit to reviewing it in July.

Lock your Anthropic SDK pricing before the pause lifts

Anthropic paused token-based billing on June 16, the day it was to take effect. Power users faced 2-3x cost increases under the proposed plan. The pause is temporary — 'working to update the plan' is not cancelled. If you use Claude Agent SDK via third-party tools (Zed, Xcode, JetBrains), check your June billing now. Enterprise contracts should lock pricing before the new plan lands.

Audit your AI deployment surface — the attack timeline just compressed

Five Eyes intelligence alliance warned that AI cyberattacks are 'months, not years' away — the same day Patch the Planet results went public. Treat unpatched LLM-facing infrastructure as critical. If you run any AI service exposed to the internet, run a security audit this week.

Track Reflection AI's first frontier model — they're now Colossus tenant #3

Reflection AI (ex-DeepMind founders) signed a $6.3B SpaceX compute deal through 2029 with no shipped frontier model. Joins Anthropic and Google as third Colossus 2 tenant. Nvidia $800M backing at $25B valuation. If you're evaluating open-weight alternatives to closed frontier models, add Reflection to your watch list — verdict pending until they ship.

Watch Cursor's vertical-integration bet — platform beyond IDE

At Cursor Ship: first in-house AI model, new Cursor Git, and a mobile app. Post-SpaceX $60B acquisition, Cursor is betting on vertical integration over ecosystem play. If you're on Cursor, watch whether this improves stability or fragments the experience. Our directory comparison tracks Cursor vs alternatives.

Evaluate Claude Tag's data-access surface before Slack adoption

Anthropic's always-on enterprise AI teammate for Slack learns organizational context from every searchable message, document, and decision. Strategic enterprise data play ahead of Anthropic's IPO. If your org uses Slack + Anthropic, evaluate what data Claude Tag can access before deployment — it becomes training surface.

Sur notre radar

  • Claude outage hits 8,000+ reports on Jun 23 — claude.ai, Claude Code, and API all affected. Agentic pipeline failures linked to sub-agent architecture. Infrastructure resilience concerns before Anthropic IPO.

  • Mythos found vulnerabilities in classified US systems — AP confirms Anthropic's model was actively finding vulns in classified systems. NSA lost access amid Trump administration dispute with Anthropic. Export ban consequences now have concrete operational cost.

  • White House differential treatment: OpenAI vs Anthropic — GPT-5.5-Cyber unrestricted while Mythos 5 banned. Raises fairness questions about Fable/Mythos export ban rationale. Axios: 'inconsistency the industry has been privately fuming about for weeks.'

  • US AI stock sell-off, Oracle lays off 21,000 — Tech stocks tumbled on AI spending sustainability concerns. Oracle's 12.9% workforce reduction explicitly attributed to AI deployment. Market signal for investment climate shift.

  • Boris Cherny: agentic loops then tempers 'AI solved coding' — Claude Code lead: loops 'as important as the step from source code to agents.' Then admits AI writing 100% of code is 'problematic' for companies. Anthropic runs 24/7 persistent agents that hunt architectural improvements and submit PRs.

  • EU AI Act Omnibus: high-risk rules delayed to Dec 2027 — Parliament voted 423-57-174. +17 months for high-risk AI rules. Nudifier/CSAM ban effective Dec 2026. GPAI timelines unchanged. Scheduling change, not deregulation — but meaningful breathing room for EU AI deployment.

  • Zapier AI switches to model-tier pricing — Standard (1x), Advanced 3x, Premium 5x. Tool calls add to base cost. 75-task per-step limit. Affects directory entity. Model-tier pricing spreading beyond API providers to platform tools.

  • Gemini 3.5 Pro GA imminent, 2M-token context — Deep Think reasoning, 2M context window (2x Fable 5's 1M). Three frontier models landing within ~10 days: Fable 5, GPT-5.6, Gemini 3.5 Pro.

Écarté

  • Consumer products: Google Home, Fitbit Air, Sony Xperia AI, Meta glasses, ByteDance Seedance 2.5 — not directory-relevant
  • Opinion/commentary: Cory Doctorow AI bubble, Sam Altman movie drama, 'Cancel Claude', How to Passive-Aggressively Shame LLM Users — no new data
  • Corporate/political: Snap Dotmo spinoff, Bernie Sanders $7T plan, AI super PACs, FERC data centers, Nvidia water use, Chevron-Microsoft power deal
  • Niche tools: darktable 5.6, Deno Desktop, GitKraken Code Flow, VoltanaLLM, various Show HN projects — early-stage, not directory-relevant
  • Academic/trend: King's study AI nuclear signaling, AI persuades humans, Chinese universities cut language majors, AI was supposed to make smarter decisions

Our take: Cette semaine a tranché un débat qui durait depuis un an : la sécurité IA n'est pas un problème futur. GPT-5.5-Cyber a trouvé 34 CVE réelles en 5 jours — validées par 25 ingénieurs de Trail of Bits avant qu'un seul correctif ne soit déployé. L'alliance de renseignement Five Eyes a averti que les attaques sont à « des mois, pas des années ». Mythos a été retiré de la NSA. La militarisation de l'IA va dans les deux sens simultanément, et l'asymétrie favorise celui qui bouge le plus vite. Pour l'instant, de justesse, ce sont les défenseurs. La vraie question n'est pas de savoir si l'IA peut casser des choses — c'est si le pipeline Patch the Planet peut tenir mois après mois. Si c'est le cas, notre verdict GPT-4o mérite un sérieux réexamen.

— Neomanex, from our own production runs

Recevez le numéro 15

Les changements de verdict et les échéances de la semaine prochaine, dans votre boîte mail.

En vous abonnant, vous acceptez notre Politique de confidentialité. Désabonnement à tout moment.