The AI startup announced this decision on its blog Friday (Aug. 7) after an in-house evaluation of its Astra model caused the company to realize it could not “rule out critical cyber capabilities” under its Preparedness Framework, which outlines how the company responds to possible AI risks.
Under the framework, an AI model reaches the “critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention,” the blog post said.
Models also reach the threshold if they can “devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal,” OpenAI added.
According to the blog post, OpenAI’s recent evaluations found that Astra made significant coding and cybersecurity progress that pushed the model closer to the “critical” threshold.
With that in mind, OpenAI says it is “pausing internal activities involving Astra that do not yet meet these strengthened security control requirements,” and that it has “implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation.”
The company also said it is working with government agencies and “select” AI safety groups to test the model’s capabilities.
The news follows a series of cybersecurity incidents involving advanced AI models. Last month, a pair of OpenAI models broke loose from their testing environment and hacked open-source AI tool provider Hugging Face.
Two days later, Anthropic said that a review of its evaluation history, prompted by OpenAI’s findings, found three incidents since April in which its Claude models had accessed the systems of three different organizations.
And Meta said last week that one of its AI models hacked another company during cybersecurity testing. The Facebook owner said an unintentional misconfiguration by a company that performs cybersecurity evaluations, Irregular, provided the model access to the internet during testing.
A report on the Astra incident by The Wall Street Journal includes comments from Jeffrey Ladish, executive director of Palisade Research, a nonprofit AI lab that examines AI capabilities to determine risks. He said OpenAI should have halted work on Astra after finding out about the Hugging Face hack.
“It’s definitely late,” he said. “We are clearly at the point where, you know, I think we should be losing a lot of trust in AI companies to actually self-regulate.”
For all PYMNTS AI coverage, subscribe to the daily AI Newsletter.