Die Woche, in der Defense-KI live ging
OpenAIs GPT-5.5-Cyber fand 34 FreeBSD-CVEs, 8 Linux-Kernel-PoCs und 5 Chrome-V8-Bugs in 5 Tagen — das erste Frontier-Modell, das bewusst als Verteidigungswaffe im Internet-Maßstab eingesetzt wurde. Die Five-Eyes-Allianz warnte am selben Tag, dass KI-Cyberangriffe Monate, nicht Jahre entfernt sind.
61 announcements scanned · 15 mattered · 0 verdicts changed
Vor der nächsten Ausgabe handeln
Re-examine our GPT-4o verdict — Patch the Planet is the reason
GPT-5.5-Cyber found 34 FreeBSD CVEs, 8 Linux kernel PoCs, and 5 Chrome V8 bugs in 5 days, validated by 25 Trail of Bits engineers. If this pipeline sustains month over month, our GPT-4o recommended call needs a hard re-examination. We commit to reviewing it in July.
Lock your Anthropic SDK pricing before the pause lifts
Anthropic paused token-based billing on June 16, the day it was to take effect. Power users faced 2-3x cost increases under the proposed plan. The pause is temporary — 'working to update the plan' is not cancelled. If you use Claude Agent SDK via third-party tools (Zed, Xcode, JetBrains), check your June billing now. Enterprise contracts should lock pricing before the new plan lands.
Audit your AI deployment surface — the attack timeline just compressed
Five Eyes intelligence alliance warned that AI cyberattacks are 'months, not years' away — the same day Patch the Planet results went public. Treat unpatched LLM-facing infrastructure as critical. If you run any AI service exposed to the internet, run a security audit this week.
Track Reflection AI's first frontier model — they're now Colossus tenant #3
Reflection AI (ex-DeepMind founders) signed a $6.3B SpaceX compute deal through 2029 with no shipped frontier model. Joins Anthropic and Google as third Colossus 2 tenant. Nvidia $800M backing at $25B valuation. If you're evaluating open-weight alternatives to closed frontier models, add Reflection to your watch list — verdict pending until they ship.
Watch Cursor's vertical-integration bet — platform beyond IDE
At Cursor Ship: first in-house AI model, new Cursor Git, and a mobile app. Post-SpaceX $60B acquisition, Cursor is betting on vertical integration over ecosystem play. If you're on Cursor, watch whether this improves stability or fragments the experience. Our directory comparison tracks Cursor vs alternatives.
Evaluate Claude Tag's data-access surface before Slack adoption
Anthropic's always-on enterprise AI teammate for Slack learns organizational context from every searchable message, document, and decision. Strategic enterprise data play ahead of Anthropic's IPO. If your org uses Slack + Anthropic, evaluate what data Claude Tag can access before deployment — it becomes training surface.
Auf unserem Radar
Claude outage hits 8,000+ reports on Jun 23 — claude.ai, Claude Code, and API all affected. Agentic pipeline failures linked to sub-agent architecture. Infrastructure resilience concerns before Anthropic IPO.
Mythos found vulnerabilities in classified US systems — AP confirms Anthropic's model was actively finding vulns in classified systems. NSA lost access amid Trump administration dispute with Anthropic. Export ban consequences now have concrete operational cost.
White House differential treatment: OpenAI vs Anthropic — GPT-5.5-Cyber unrestricted while Mythos 5 banned. Raises fairness questions about Fable/Mythos export ban rationale. Axios: 'inconsistency the industry has been privately fuming about for weeks.'
US AI stock sell-off, Oracle lays off 21,000 — Tech stocks tumbled on AI spending sustainability concerns. Oracle's 12.9% workforce reduction explicitly attributed to AI deployment. Market signal for investment climate shift.
Boris Cherny: agentic loops then tempers 'AI solved coding' — Claude Code lead: loops 'as important as the step from source code to agents.' Then admits AI writing 100% of code is 'problematic' for companies. Anthropic runs 24/7 persistent agents that hunt architectural improvements and submit PRs.
EU AI Act Omnibus: high-risk rules delayed to Dec 2027 — Parliament voted 423-57-174. +17 months for high-risk AI rules. Nudifier/CSAM ban effective Dec 2026. GPAI timelines unchanged. Scheduling change, not deregulation — but meaningful breathing room for EU AI deployment.
Zapier AI switches to model-tier pricing — Standard (1x), Advanced 3x, Premium 5x. Tool calls add to base cost. 75-task per-step limit. Affects directory entity. Model-tier pricing spreading beyond API providers to platform tools.
Gemini 3.5 Pro GA imminent, 2M-token context — Deep Think reasoning, 2M context window (2x Fable 5's 1M). Three frontier models landing within ~10 days: Fable 5, GPT-5.6, Gemini 3.5 Pro.
Aussortiert
- Consumer products: Google Home, Fitbit Air, Sony Xperia AI, Meta glasses, ByteDance Seedance 2.5 — not directory-relevant
- Opinion/commentary: Cory Doctorow AI bubble, Sam Altman movie drama, 'Cancel Claude', How to Passive-Aggressively Shame LLM Users — no new data
- Corporate/political: Snap Dotmo spinoff, Bernie Sanders $7T plan, AI super PACs, FERC data centers, Nvidia water use, Chevron-Microsoft power deal
- Niche tools: darktable 5.6, Deno Desktop, GitKraken Code Flow, VoltanaLLM, various Show HN projects — early-stage, not directory-relevant
- Academic/trend: King's study AI nuclear signaling, AI persuades humans, Chinese universities cut language majors, AI was supposed to make smarter decisions
Our take: Diese Woche hat eine einjährige Debatte beendet: KI-Sicherheit ist kein Zukunftsproblem. GPT-5.5-Cyber fand 34 echte CVEs in 5 Tagen — validiert von 25 Trail of Bits-Ingenieuren, bevor ein einziger Patch ausgeliefert wurde. Die Five-Eyes-Geheimdienstallianz warnte, dass Angriffe "Monate, nicht Jahre" entfernt sind. Mythos wurde der NSA entzogen. Die Bewaffnung von KI verläuft gleichzeitig in beide Richtungen, und die Asymmetrie begünstigt, wer sich schneller bewegt. Im Moment sind das — knapp — die Verteidiger. Die eigentliche Frage ist nicht, ob KI Dinge kaputtmachen kann — sondern ob die Patch-the-Planet-Pipeline Monat für Monat durchhält. Wenn ja, verdient unser GPT-4o-Urteil eine gründliche Neubewertung.
— Neomanex, from our own production runs
Ausgabe 15 erhalten
Die Urteilsänderungen und Fristen der nächsten Woche, in deinem Postfach.