La semana en que la IA de defensa entró en acción
El GPT-5.5-Cyber de OpenAI encontró 34 CVEs de FreeBSD, 8 PoCs del kernel de Linux y 5 errores de Chrome V8 en 5 días — el primer modelo frontera desplegado deliberadamente como arma de defensa a escala de internet. La alianza Five Eyes advirtió el mismo día que los ciberataques con IA están a meses, no años, de distancia.
61 announcements scanned · 15 mattered · 0 verdicts changed
Actúa antes del próximo número
Re-examine our GPT-4o verdict — Patch the Planet is the reason
GPT-5.5-Cyber found 34 FreeBSD CVEs, 8 Linux kernel PoCs, and 5 Chrome V8 bugs in 5 days, validated by 25 Trail of Bits engineers. If this pipeline sustains month over month, our GPT-4o recommended call needs a hard re-examination. We commit to reviewing it in July.
Lock your Anthropic SDK pricing before the pause lifts
Anthropic paused token-based billing on June 16, the day it was to take effect. Power users faced 2-3x cost increases under the proposed plan. The pause is temporary — 'working to update the plan' is not cancelled. If you use Claude Agent SDK via third-party tools (Zed, Xcode, JetBrains), check your June billing now. Enterprise contracts should lock pricing before the new plan lands.
Audit your AI deployment surface — the attack timeline just compressed
Five Eyes intelligence alliance warned that AI cyberattacks are 'months, not years' away — the same day Patch the Planet results went public. Treat unpatched LLM-facing infrastructure as critical. If you run any AI service exposed to the internet, run a security audit this week.
Track Reflection AI's first frontier model — they're now Colossus tenant #3
Reflection AI (ex-DeepMind founders) signed a $6.3B SpaceX compute deal through 2029 with no shipped frontier model. Joins Anthropic and Google as third Colossus 2 tenant. Nvidia $800M backing at $25B valuation. If you're evaluating open-weight alternatives to closed frontier models, add Reflection to your watch list — verdict pending until they ship.
Watch Cursor's vertical-integration bet — platform beyond IDE
At Cursor Ship: first in-house AI model, new Cursor Git, and a mobile app. Post-SpaceX $60B acquisition, Cursor is betting on vertical integration over ecosystem play. If you're on Cursor, watch whether this improves stability or fragments the experience. Our directory comparison tracks Cursor vs alternatives.
Evaluate Claude Tag's data-access surface before Slack adoption
Anthropic's always-on enterprise AI teammate for Slack learns organizational context from every searchable message, document, and decision. Strategic enterprise data play ahead of Anthropic's IPO. If your org uses Slack + Anthropic, evaluate what data Claude Tag can access before deployment — it becomes training surface.
En nuestro radar
Claude outage hits 8,000+ reports on Jun 23 — claude.ai, Claude Code, and API all affected. Agentic pipeline failures linked to sub-agent architecture. Infrastructure resilience concerns before Anthropic IPO.
Mythos found vulnerabilities in classified US systems — AP confirms Anthropic's model was actively finding vulns in classified systems. NSA lost access amid Trump administration dispute with Anthropic. Export ban consequences now have concrete operational cost.
White House differential treatment: OpenAI vs Anthropic — GPT-5.5-Cyber unrestricted while Mythos 5 banned. Raises fairness questions about Fable/Mythos export ban rationale. Axios: 'inconsistency the industry has been privately fuming about for weeks.'
US AI stock sell-off, Oracle lays off 21,000 — Tech stocks tumbled on AI spending sustainability concerns. Oracle's 12.9% workforce reduction explicitly attributed to AI deployment. Market signal for investment climate shift.
Boris Cherny: agentic loops then tempers 'AI solved coding' — Claude Code lead: loops 'as important as the step from source code to agents.' Then admits AI writing 100% of code is 'problematic' for companies. Anthropic runs 24/7 persistent agents that hunt architectural improvements and submit PRs.
EU AI Act Omnibus: high-risk rules delayed to Dec 2027 — Parliament voted 423-57-174. +17 months for high-risk AI rules. Nudifier/CSAM ban effective Dec 2026. GPAI timelines unchanged. Scheduling change, not deregulation — but meaningful breathing room for EU AI deployment.
Zapier AI switches to model-tier pricing — Standard (1x), Advanced 3x, Premium 5x. Tool calls add to base cost. 75-task per-step limit. Affects directory entity. Model-tier pricing spreading beyond API providers to platform tools.
Gemini 3.5 Pro GA imminent, 2M-token context — Deep Think reasoning, 2M context window (2x Fable 5's 1M). Three frontier models landing within ~10 days: Fable 5, GPT-5.6, Gemini 3.5 Pro.
Descartado
- Consumer products: Google Home, Fitbit Air, Sony Xperia AI, Meta glasses, ByteDance Seedance 2.5 — not directory-relevant
- Opinion/commentary: Cory Doctorow AI bubble, Sam Altman movie drama, 'Cancel Claude', How to Passive-Aggressively Shame LLM Users — no new data
- Corporate/political: Snap Dotmo spinoff, Bernie Sanders $7T plan, AI super PACs, FERC data centers, Nvidia water use, Chevron-Microsoft power deal
- Niche tools: darktable 5.6, Deno Desktop, GitKraken Code Flow, VoltanaLLM, various Show HN projects — early-stage, not directory-relevant
- Academic/trend: King's study AI nuclear signaling, AI persuades humans, Chinese universities cut language majors, AI was supposed to make smarter decisions
Our take: Esta semana resolvió un debate de un año: la seguridad con IA no es un problema futuro. GPT-5.5-Cyber encontró 34 CVEs reales en 5 días — validados por 25 ingenieros de Trail of Bits antes de que se enviara un solo parche. La alianza de inteligencia Five Eyes advirtió que los ataques están a 'meses, no años' de distancia. Mythos fue retirado de la NSA. La militarización de la IA va en ambas direcciones simultáneamente, y la asimetría favorece a quien se mueve más rápido. Ahora mismo, por poco, esos son los defensores. La verdadera pregunta no es si la IA puede romper cosas — es si el pipeline de Patch the Planet puede sostenerse mes tras mes. Si puede, nuestro veredicto de GPT-4o merece un reexamen profundo.
— Neomanex, from our own production runs
Recibe el número 15
Los cambios de veredicto y fechas clave de la próxima semana, en tu bandeja de entrada.