Artificial intelligence (AI) labs and cybersecurity firms are considering ways to give advanced AI models controlled access to the internet during testing, rather than trying to keep the models confined to a sandbox, Bloomberg reported Tuesday (Aug. 25).
This consideration comes after OpenAI, Anthropic and Meta disclosed separate incidents in which models escaped a sandbox, accessed the internet and breached other companies’ real servers, according to the report.
Giving advanced models controlled access to the internet would allow companies to get a better understanding of the models’ true capabilities, advocates of the change say. However, this move would also give the models access to systems that are outside the intended testing, per the report.
We’d love to be your preferred source for news.
Please add us to your preferred sources list so our news, data and interviews show up in your feed. Thanks!
In one of the reported incidents in which models reached the internet during testing, OpenAI said July 21 that a combination of its models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production database in search of a solution to the evaluation problem.
OpenAI said at the time that the company considered the incident to be “an unprecedented cyber incident, involving state-of-the-art cyber capabilities.”
It was reported July 31 that while examining this hacking incident, OpenAI unearthed other examples of its autonomous AI agents breaking containment.
By Aug. 5, it was reported that Anthropic and Meta had experienced separate incidents in which AI models hacked another company during cybersecurity testing.
According to the report, while the incident at OpenAI saw an AI agent independently exploit a previously unknown vulnerability to reach the internet during testing, the incidents at Anthropic and Meta resulted from unintentional misconfigurations by the company that conducted the cybersecurity evaluations.
The Wall Street Journal said of the Meta incident: “The new case is the latest proof that AI loss-of-control scenarios, once confined to science fiction and AI-safety experiments, are now a real-world issue.”
OpenAI said Aug. 7 that it was pausing internal activity related to a new AI model due to security concerns. The company said that an in-house evaluation of its Astra model caused OpenAI to realize it could not “rule out critical cyber capabilities” under its Preparedness Framework, which outlines how the company responds to possible AI risks.