OpenAI says Astra can find zero-day flaws: capabilities and safeguards explained
OpenAI says its forthcoming Astra model has reached the “Critical” cybersecurity capability threshold under the company’s Preparedness Framework, the first OpenAI model to receive that designation. The assessment means the company believes Astra can, with suitable tools and access, find previously unknown vulnerabilities and develop working exploits across hardened systems with less continuous human direction.
This is a company assessment released before the model’s full launch, not evidence that Astra has been made broadly available for unrestricted offensive use. OpenAI says access to the most advanced cybersecurity functions will initially be limited.
What OpenAI’s Critical threshold means
The framework describes two routes to the threshold: independently identifying and developing functional zero-day exploits across many hardened real-world systems, or planning and executing novel end-to-end attacks against hardened targets from a high-level objective.
OpenAI reports that Astra scored 100% on ExploitBench, a benchmark based on known vulnerabilities. Because public benchmarks can be contaminated by training data, the company also assembled an internal test using 20 recently disclosed high-severity V8 vulnerabilities. It says Astra achieved higher code-execution rates than GPT-5.6 Sol while using fewer output tokens and found two previously unknown vulnerabilities during one exploit chain. OpenAI says it is disclosing those flaws to maintainers.
In expert-led evaluations, the company says Astra built a browser compromise that escaped a sandbox and executed commands on the host, and combined operating-system weaknesses into a local privilege-escalation chain. Full details have not been released, partly because disclosure could create risk before fixes are available.
Why the company delayed parts of development
OpenAI says it paused or delayed some Astra training and release work while strengthening network isolation, monitoring, misuse controls and alignment testing. The company restarted one larger reinforcement-learning run on August 28 after introducing additional requirements, while continuing to hold back some experimental runs.
Safeguards described for release
- Additional training to refuse disallowed cyber requests.
- System-level classifiers and cross-conversation misuse monitoring.
- Monitoring designed to stop potentially unauthorised model actions.
- Restricted access for advanced cyber work, initially through testers and later defensive programmes such as Daybreak Blue.
OpenAI reports a 91.5% refusal rate on its cyber-jailbreak evaluation for Astra, compared with 59% for GPT-5.6 Sol. These are OpenAI’s own evaluation results and should be interpreted alongside the full system card when it is released.
What users may notice
The company acknowledges that stronger monitoring may slow, pause or stop legitimate long-running work, including defensive security tasks. ChatGPT or Codex users could be asked to review an action; API work may stop when a monitor intervenes.
The central issue is not simply whether an AI model can find vulnerabilities. It is whether access control, auditing, isolation and human authorisation can keep those capabilities directed toward defence. Independent scrutiny will be important after Astra’s system card and production behaviour are available.
Source: OpenAI’s Path to Astra assessment.
Official artwork accompanying the September 1 assessment. Image: OpenAI.

