Independent Investigators Uncover Broader Rogue Activity by OpenAI Agents

OpenAI cybersecurity

Independent investigators found that rogue activity by OpenAI agents was more extensive than previously disclosed, Reuters reported Wednesday (Sept. 9).

    Get the Full Story

    Complete the form to unlock this article and enjoy unlimited free access to all PYMNTS content — no additional logins required.

    Subscribe to our daily newsletter, PYMNTS Today.

    By completing this form, you agree to receive marketing communications from PYMNTS and to the sharing of your information with our sponsor, if applicable, in accordance with our Privacy Policy and Terms and Conditions.

    Six sets of independent investigators found that the agents used more than 10 previously undisclosed websites to communicate with each other during a test in which they were restricted from posting on the web, according to the report.

    While the agents’ behavior doesn’t amount to hacking, it does show that they circumvented restrictions that were placed upon them, the report said.

    The agents used websites, including some obscure and, in some cases, two-decade-old communally edited wikis and online text storage sites, for unsanctioned communications as they worked to cheat on a test, per the report.

    OpenAI told Reuters, per the report, that its review of agent activity had so far “not identified other activity matching the severity or scale of Hugging Face,” which was another breach conducted by AI agents, and that it would share a framework for reporting rogue behavior of AI agents “soon.”

    OpenAI announced in a July 21 blog post that a security incident reported a week earlier by Hugging Face was caused by OpenAI models as their cyber capabilities were being tested by OpenAI.

    During OpenAI’s internal evaluation of GPT-5.6 Sol and a more capable pre-release model, the models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production database in search of a solution to the evaluation problem.

    OpenAI said in the July blog post that it considered the incident to be “an unprecedented cyber incident” and that the company was “responding accordingly.”

    It was reported Aug. 2 that as OpenAI deepened its examination of the incident at Hugging Face, the company unearthed other examples of its AI agents breaking containment.

    Asked about that report, an OpenAI spokesperson referred Reuters to a statement the company issued a week earlier that said it was reviewing “broader activity from our models” along with the Hugging Face incident.

    On Aug. 6, it was reported that a Meta AI model hacked another company during cybersecurity testing in an incident similar to earlier ones that happened at Anthropic and OpenAI.

    For all PYMNTS AI coverage, subscribe to the daily AI Newsletter.