OpenAI paused development of the model Astra after reaching the critical cybersecurity threshold
OpenAI paused part of the development of the forthcoming model Astra after an internal review under Preparedness Framework showed that the model had reached the critical cybersecurity threshold and can independently develop zero-day exploits. The model is unrelated to the attack on Hugging Face.
On 7 August 2026, OpenAI announced that it had paused development of some parts of the forthcoming model Astra. According to an internal review, the model demonstrated significant progress in agentic coding and cybersecurity and reached the so-called critical cybersecurity threshold under Preparedness Framework – a tool that OpenAI created in 2023 to assess the capabilities of the most advanced models.
Under the terms of this framework, a model reaches the critical threshold if it can, without human intervention, identify and develop working zero-day exploits of any severity in many secure, real-world critical systems, or if it can devise and execute new cyberattack strategies from start to finish against well-protected targets when given only a broadly defined objective. OpenAI stated that evaluation of the model Astra is still ongoing, but the preliminary results are strong enough that a critical level of capability cannot be ruled out at this point. According to the company, the model Astra was unrelated to the recent attack on Hugging Face.
In response, OpenAI is introducing stricter security controls for models with greater capabilities and comprehensive monitoring of risky actions across all agentic applications using the model Astra, and has paused internal activities involving this model that do not meet the new security requirements. The company also stated that it will work with government agencies and selected AI safety organizations to test the capabilities of the model and that it is providing recommended security controls to external testing partners. The company explained the disclosure by saying that it considers it important to be transparent with the public and the security community about a possible shift in model capabilities.
Why it matters
This is a public acknowledgment by OpenAI that one of its models has reached a threshold beyond which it could independently find and exploit zero-day vulnerabilities in protected systems – a capability previously associated more with specialized attackers or defense teams. For security and compliance teams, this is a concrete signal that agentic models are approaching a level that requires specific control mechanisms, monitoring and cooperation with external testing partners, rather than just general statements about safety.
What was added since the original report
Verified updates
-
The model identifies and develops working zero-day exploits of all severity levels without human intervention; The model is not linked to the recent attacks on Hugging Face; OpenAI is introducing stricter security controls and monitoring
- The model identifies and develops working zero-day exploits of all severity levels without human intervention
- The model is not linked to the recent attacks on Hugging Face
- OpenAI is introducing stricter security controls and monitoring
Relevant practical impact
What this means
For a business
Companies deploying agentic AI tools or working with OpenAI on model testing can expect stricter security controls, comprehensive monitoring of risky actions and recommended security controls for testing partners, which places greater demands on their own security and compliance processes.
Risks and complianceCheck the original
Event sources
confirmed by 2 independent sources · 2 publishers, 2 independent. We count feeds from the same owner only once.