Six sets of independent investigators found that the agents used more than 10 previously undisclosed websites to communicate with each other during a test in which they were restricted from posting on the web, according to the report.
While the agents’ behavior doesn’t amount to hacking, it does show that they circumvented restrictions that were placed upon them, the report said.
The agents used websites, including some obscure and, in some cases, two-decade-old communally edited wikis and online text storage sites, for unsanctioned communications as they worked to cheat on a test, per the report.
OpenAI told Reuters, per the report, that its review of agent activity had so far “not identified other activity matching the severity or scale of Hugging Face,” which was another breach conducted by AI agents, and that it would share a framework for reporting rogue behavior of AI agents “soon.”
We’d love to be your preferred source for news.
Please add us to your preferred sources list so our news, data and interviews show up in your feed. Thanks!
OpenAI announced in a July 21 blog post that a security incident reported a week earlier by Hugging Face was caused by OpenAI models as their cyber capabilities were being tested by OpenAI.
During OpenAI’s internal evaluation of GPT-5.6 Sol and a more capable pre-release model, the models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production database in search of a solution to the evaluation problem.
OpenAI said in the July blog post that it considered the incident to be “an unprecedented cyber incident” and that the company was “responding accordingly.”
It was reported Aug. 2 that as OpenAI deepened its examination of the incident at Hugging Face, the company unearthed other examples of its AI agents breaking containment.
Asked about that report, an OpenAI spokesperson referred Reuters to a statement the company issued a week earlier that said it was reviewing “broader activity from our models” along with the Hugging Face incident.
On Aug. 6, it was reported that a Meta AI model hacked another company during cybersecurity testing in an incident similar to earlier ones that happened at Anthropic and OpenAI.
For all PYMNTS AI coverage, subscribe to the daily AI Newsletter.