Astra Becomes OpenAI's First Model to Reach 'Critical' Cybersecurity Level

OpenAI announced on Tuesday (2) that Astra is the company’s first model to achieve 'Critical' status under its Preparedness Framework. 100% on ExploitBench and two zero-days discovered independently triggered access gating.
On September 2, OpenAI announced that Astra is the first model from the company to cross the 'Critical' threshold in cybersecurity capabilities according to the Preparedness Framework, the internal risk framework the company has adopted since 2023. This classification is reserved for systems capable of identifying and exploiting zero-day vulnerabilities in hardened targets without step-by-step human guidance, coinciding with a business decision: Astra will go into production soon, but the more sensitive functions will be restricted to a small group of partners through the Daybreak Blue program.
The classification is based on a battery of internal tests that OpenAI detailed. In the ExploitBench, a benchmark for developing exploits from known vulnerabilities, the model scored 100%. In a second test, involving 20 high-severity vulnerabilities disclosed between June and August of this year, Astra discovered and exploited two zero-days as part of a complete exploitation chain. In cybersecurity jailbreaking assessments, it rejected 91.5% of malicious requests, compared to 59% for GPT-5.6 Sol in the same set.
The company halted development in August to add layers of monitoring, containment controls, and additional training. According to communications from OpenAI, Astra now includes jailbreak detectors, sandbox escape assessments, and reasoning chain monitoring. Governments and external security groups will be involved in upcoming tests, signaling that the company aims to secure institutional validation before broad release.
Why 'Critical' Changes the Conversation
The label is not commercial; it’s operational. For a CISO already operating under the U.S. Executive Order on AI and the European AI Act (both require evaluations before frontier systems are put into production), the arrival of a model classified at the highest level by its own provider anticipates the regulator's roadmap. The U.S. Department of Commerce has maintained early access agreements with OpenAI and Anthropic since 2024, and the European AI Office now has grounds to demand equivalent assessments under Article 51 of the AI Act, whose second wave of obligations took effect on August 2.
Access through Daybreak Blue replicates a pattern that Anthropic has been using with Claude Mythos: offensive capabilities remain with a verified subset, while defensive capabilities go to market. This is a practical response to the risk of asymmetry, but it creates a new commercial asymmetry. Active contract red team firms will now have access to frontier tools that smaller competitors will not.
Where the Effect Will Be Felt
Two markets will feel the impact soon. In the United States, providers like Mandiant, CrowdStrike, and Palo Alto Networks themselves should immediately position themselves within Daybreak Blue to avoid losing ground in managed pentesting; federal contracts under the Cybersecurity National Action Plan already reward operators that bring models with certified red teams. In Germany, BSI has signaled that any offensive analysis product used in critical infrastructure will require origin auditing, which works against short rollout timelines and shifts risk to local providers like T-Systems.
Unit 42, Palo Alto Networks' incident response arm, published data on the same day that underscores the concern: the fastest 25% of incidents that the firm investigated in 2025 reached exfiltration in 72 minutes, compared to 285 minutes the previous year. If a 'Critical' model becomes available to an adversary with average operational competence, the window for detection narrows even further.
The Counterpoint That Needs to Stand
Analysts who track AI security benchmarks remind us that ExploitBench and derived internal tests are manufactured by the same lab and do not replace controlled exposure in live environments. The 100% performance on a set designed by OpenAI serves both as an argument for the classification and a natural target for scrutiny: no board will feel comfortable with such metrics until external testers publish replications. This is why Anthropic began releasing its red team reports prior to launch, and Astra should be held to the same standard.
What to Watch for in the Coming Weeks
Three milestones are important. First, whether the European AI Office will require an assessment under Article 51 before release in the EU, which could delay regional distribution. Second, which names will enter the Daybreak Blue list: any notable absence of an Asian Big Tech company would be a geopolitical, not just commercial, issue. Third, whether Anthropic responds by publishing an equivalent classification of Mythos 5.1, which would force a side-by-side comparison that the market lacks today.