Now, a pair of reports—one by OpenAI and the other from independent researchers—offers new details into the cybersecurity incident.
For example, independent investigators METR and Redwood Research, which had been brought in by OpenAI, found that the breach wasn’t the result of one rogue artificial intelligence (AI) agent but a “swarm” of around 700 of them.
Meanwhile, OpenAI found examples of its agents trying to “cheat” at tasks by finding solutions to their given problems online.
“This behavior is known as ‘reward hacking,’ in which a model finds an unintended way to achieve an outcome that earns reward without completing the task in the way the evaluation was designed to measure,” the report said.
The two reports said that AI models tried to hide their misbehavior by attempting to delete or modify messages that would provide records of their actions.
We’d love to be your preferred source for news.
Please add us to your preferred sources list so our news, data and interviews show up in your feed. Thanks!
Jeffrey Ladish, whose Palisade Research studies AI agents, told NBC News that cheating on non-cyber tests indicates misbehavior that could be more deeply rooted.
“It’s sort of like asking, ‘If Billy cheats in every class instead of just computer class, is that more concerning?’ And the answer is, well, ‘Yes it’s more concerning,’” he said.
Hugging Face, whose platform hosts AI datasets, reported an AI-powered data breach of its site in mid-July. The company said a dataset uploaded to its platform exploited a security weakness to run malicious code on its servers, allowing hackers to escalate their permissions and get greater access to the company’s internal systems.
A few days later, OpenAI said the incident had been caused by its models as the startup was testing their capabilities.
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” OpenAI wrote on its blog.
Incidents such as these have led regulators to take a closer look at the controls companies are using to keep AI models in check.
For example, Alabama Attorney General Steve Marshall announced an investigation into OpenAI earlier this week related to the Hugging Face incident.
Marshall has subpoenaed OpenAI as part of a probe into what his office deemed “the company’s complete lack of oversight and adequate safeguards.”