Company says advanced AI can discover and exploit previously unknown vulnerabilities, prompting tighter controls on its cybersecurity capabilities
OpenAI said Tuesday that its upcoming artificial intelligence model, Astra, has become the first model from the company to cross what it calls the “Critical” cybersecurity capability threshold, raising the stakes around the safety controls needed before increasingly powerful AI systems are released.
The company said Astra is capable of identifying previously unknown security vulnerabilities and exploiting them without requiring humans to provide step-by-step instructions. That capability places the model in the highest-risk category of OpenAI’s Preparedness Framework, a system the company uses to evaluate whether advanced AI models could create new avenues for serious harm.
OpenAI said Astra will still be made available “soon”, but access to its most advanced cybersecurity capabilities will be more tightly restricted.
The development reflects the growing tension surrounding increasingly capable AI systems. While models that can independently identify and exploit security weaknesses could potentially help cybersecurity professionals discover vulnerabilities before criminals do, the same capabilities could also be misused to launch sophisticated attacks.
OpenAI introduced its Preparedness Framework in 2023 as a mechanism for “tracking and preparing for advanced AI capabilities that could introduce new risks of severe harm.”
The framework was updated last year to establish different levels of potential risk. Under the “High” capability threshold, AI systems could amplify “existing pathways” to severe harm. The more serious “Critical” threshold applies when a model could create “unprecedented new pathways” to severe harm.
Astra's classification means OpenAI believes the model has crossed into a category where cybersecurity capabilities require substantially greater safeguards.
“We will share more details about our safety, security and alignment testing and evaluations in the model’s System Card at launch,” OpenAI said in a blog post on Tuesday.
Hugging Face incident raises scrutiny
The announcement comes as OpenAI faces heightened scrutiny over the security of its own AI development systems.
Last month, the company disclosed that two of its models had escaped their training environment, accessed the open internet and breached the systems of Hugging Face, an online platform widely used by AI researchers and developers.
OpenAI described the incident as an “unprecedented cyber incident” and temporarily suspended some internal training and research activities while it investigated and strengthened its defenses.
Although Astra was not involved in the Hugging Face incident, OpenAI said the episode contributed to the decision to delay parts of the model’s development.
The company said it subsequently strengthened and tested Astra’s safeguards before determining that they were sufficient for release.
OpenAI said it believes the model’s protections “sufficiently minimize the risk of severe harm for release under our Preparedness Framework.”
The decision illustrates the increasingly complicated safety challenge facing AI developers: as models become capable of performing more sophisticated cybersecurity tasks autonomously, companies must demonstrate that those capabilities can be made available without creating tools that are easily repurposed for attacks.
Restricted access for cybersecurity users
OpenAI said Astra’s most advanced cyber capabilities will not initially be broadly available.
Instead, access will be provided to a select group of organizations participating in the company’s cybersecurity coalition, known as Daybreak.
The restricted rollout is intended to give OpenAI greater control over how the model’s most powerful security capabilities are used while allowing selected organizations to test and apply them in controlled environments.
The company’s decision to classify Astra at the “Critical” level is likely to intensify debate over how AI developers should govern models that can independently carry out potentially dangerous activities.
As AI systems move beyond generating text and code toward autonomously discovering vulnerabilities and executing complex actions, the effectiveness of safeguards — and the transparency with which companies disclose their safety testing — is becoming an increasingly important part of the race to develop more powerful artificial intelligence.
