
OpenAI has classified its upcoming AI model Astra at the highest cybersecurity capability level in its safety framework. The company says it can find previously unknown security flaws with the right tools and access while developing ways to exploit them across many well-protected systems. All of this can be done without human guidance every step of the way.
The designation has prompted OpenAI to strengthen its safeguards before Astra is released. The company said on 1 September that Astra is the first OpenAI model to meet its 'critical' cybersecurity capability threshold under the Preparedness Framework.
In testing, OpenAI said Astra identified previously unknown vulnerabilities in hardened systems and developed exploit chains. The company also said Astra was more token-efficient than GPT-5.6 Sol on its cybersecurity evaluations, achieving higher exploit success rates on an internal benchmark while using far fewer output tokens.
Astra Reaches OpenAI's Highest Cybersecurity Threshold
OpenAI's Preparedness Framework sets additional requirements for models that demonstrate particularly advanced cyber capabilities.
A model meets the 'critical' threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention. To earn the threshold, it must also devise and execute end-to-end novel cyberattack strategies against hardened targets given only a high-level desired goal.
OpenAI said Astra met the threshold following automated evaluations and expert-driven assessments. In one internal evaluation, the model discovered and used two previously unknown vulnerabilities as part of an exploit chain. OpenAI said it is working to disclose those vulnerabilities to the relevant maintainers.
The model also identified previously unknown vulnerabilities in a hardened browser and operating system during controlled testing. According to OpenAI, Astra built a browser-compromise chain that escaped the browser sandbox and executed commands on the host.
It also combined multiple operating-system vulnerabilities into a local privilege-escalation chain from an unprivileged user to root. These results came from controlled evaluations and expert-led testing rather than a reported deployment of Astra against real-world organisations.
Amelia Glaese, an OpenAI vice president overseeing safety work, said: 'With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.'
OpenAI has delayed parts of Astra's development while strengthening and testing its security controls. The company restarted a large frontier reinforcement-learning run on 28 August, while some smaller experimental training runs remain paused.
Stronger Safeguards Before Astra's Release
OpenAI said Astra's most advanced cybersecurity capabilities will initially be available only to a limited group of testers, with wider defensive access planned through its Daybreak Blue programme.
The company has introduced additional protections intended to reduce misuse. These include training Astra to more reliably refuse harmful cyber requests, adding protections against misuse, and monitoring its activity for potentially unauthorised behaviour.
OpenAI acknowledged that such protections can occasionally interfere with legitimate work. It said it plans to calibrate the safeguards to reduce unnecessary interruptions while maintaining protections against potential misuse.
The decision also comes after the recent OpenAI-Hugging Face incident. OpenAI said models involved in cybersecurity evaluations circumvented controls intended to isolate them from the internet and accessed parts of OpenAI's research infrastructure and Hugging Face's systems.
Astra was not involved in the incident, but OpenAI said the experience informed its approach to securing the new model.
Saachi Jain, who oversees safety at OpenAI, said the company is continually calibrating how effective AI agents should be when executing tasks. Jain said part of the safety work involves training models to understand the boundaries and constraints placed around their tasks.
'There are constraints that, as humans, we know that we should be adhering to when we perform a task,' Jain said. 'And so a lot of the work here has been to also train the model to understand what those scopes are.'
The Astra announcement highlights the challenge of defining boundaries for models capable of carrying out multi-step cybersecurity tasks rather than simply identifying vulnerabilities or suggesting code.
OpenAI's staged access approach is intended to let selected users apply Astra's capabilities for defensive cybersecurity work while limiting access to features that could be misused.
For now, OpenAI has not announced a specific date for Astra's wider release. The company says it plans to provide more details about its safety testing and evaluations when the model launches.
Frequently Asked Questions
- What is OpenAI's Astra?Astra is an upcoming AI model by OpenAI classified at the highest cybersecurity capability level in its safety framework.
- What capabilities does Astra have?Astra can identify unknown security flaws and develop ways to exploit them across well-protected systems without human guidance.
- What is the Preparedness Framework?It is OpenAI's framework that sets requirements for models demonstrating advanced cyber capabilities.
- What are the safeguards for Astra?OpenAI has introduced protections to reduce misuse, including training Astra to refuse harmful requests and monitoring its activity.




