OpenAI Formally Designates Astra 'Critical' — First Model at the Level

OpenAI logoOpenAIFYI3 septembre 2026Modèles
Ce qui s’est passé
OpenAI formally designated Astra as meeting the Critical cybersecurity capability threshold — the first model at that level — with ExploitBench 100% and two zero-days found during evaluation.
Pourquoi c’est important
Astra stays pending in the directory (no GA date or pricing), but the designation plus a 'launch soon' path sets the template for gated frontier cyber access via alpha testers and Daybreak Blue.
Que faire
Watch Astra's directory entry for a release-timeline signal; security teams should track Daybreak Blue expansion as the access route.

OpenAI has crossed its own Rubicon. In its September 1 "Path to Astra" post, the company formally designates Astra as meeting the Critical cybersecurity capability threshold under its Preparedness Framework — the first model OpenAI has placed at that level (OpenAI, 2026). Our directory verdict stays pending: Astra still has no launch date or pricing, and OpenAI says it will publish the full system card at launch. But the governance fact is now official, not preliminary — and the gated access model OpenAI is attaching to the release is the real template to watch.

What happened

Astra's designation upgrades the August disclosure from "might reach critical" to "we now believe it meets the threshold," backed by new evidence (OpenAI, 2026):

  • ExploitBench: a perfect 100% score on the benchmark for developing exploits from known vulnerabilities.
  • Internal V8 port: on 20 high-severity V8 vulnerabilities disclosed June–August, Astra achieves much higher arbitrary-code-execution rates than GPT-5.6 Sol using far fewer output tokens — and discovered two zero-days during evaluation, now in disclosure to maintainers.
  • Expert-led tests: built a full browser-compromise chain that escaped the sandbox and executed host commands, plus a local privilege-escalation-to-root chain on a hardened OS.
  • Jailbreak resistance: refuses 91.5% of cyber-jailbreak requests versus 59% for GPT-5.6 Sol; in honeypot tests Astra never attempted to compromise surrounding infrastructure where GPT-5.6 Sol (without safeguards) did in 56% of tests.
  • Alignment: OpenAI calls it "our most aligned model to date," with chain-of-thought monitoring deployed for Astra-class models.

Two caveats keep the headline honest. OpenAI states the Astra figures reflect capabilities with Daybreak Blue access, not the default production configuration — read the benchmark bullets as "best-case with access," not "what ships by default." And the honeypot numbers describe test conditions without cyber safeguards, not normal production use.

The large frontier RL run paused after the Hugging Face incident restarted August 28 under new safety and security requirements (OpenAI, 2026). Astra was not the Hugging Face model, but OpenAI says retrospective testing shows its then-current production safeguards would have prevented that incident, and it has added stronger protections since.

Access is deliberately narrow: advanced cybersecurity work goes first to a small group of alpha testers, then expands through Daybreak Blue for defensive use. OpenAI warns the extra checks will occasionally slow, pause, or stop legitimate work.

Why it matters

Astra being formally "Critical" is the first time a frontier lab has classified one of its own models at the top rung of its safety framework — with a release path attached. The designation is the story for anyone tracking how gated cyber capability is being productized: ExploitBench-perfect models are no longer hypothetical, and the access model (alpha testers → Daybreak Blue) is now the template OpenAI is using for frontier cyber.

What changes for you

  • No change to your setup yet. Astra is not GA, has no pricing, and advanced cyber access is restricted. Its directory entry remains pending until OpenAI ships it.
  • Security teams: the Daybreak Blue expansion path matters more than the model itself — it is the only route to frontier cyber capability without a custom government program.
  • Everyone else: expect false-positive slowdowns if you use Astra-class models in ChatGPT or Codex once they ship — OpenAI says legitimate agent runs will occasionally be paused for review.

Que faire

  1. 1 Watch Astra's directory entry for a release-timeline signal — no GA date is set.
  2. 2 Security teams: track Daybreak Blue expansion as the practical access route to Critical-class cyber capability.

Outils et modèles concernés

Ne ratez plus jamais une mise à jour

Le récap hebdomadaire — uniquement les changements de verdict et les actions urgentes. Sans remplissage.

En vous abonnant, vous acceptez notre Politique de confidentialité. Désabonnement à tout moment.