The incident occurred in January but was not discovered during Anthropic’s earlier scan of transcripts that uncovered three other incidents that the company disclosed on July 30, the company said in the post.
“All four incidents occurred during cybersecurity evaluations built by the same evaluation partner,” Anthropic said in the post. “Claude was told it was operating in a simulation without internet access, but, due to a misconfiguration, it was mistakenly connected to the open internet. As is standard for cybersecurity evaluations, the models ran without the cyber safeguards that ship with our released models.”
When disclosing the three earlier incidents in a July 30 announcement, Anthropic said it found the incidents while reviewing its own cybersecurity evaluations after OpenAI disclosed that several of its models had broken out of an isolated test environment and accessed the production infrastructure of Hugging Face.
In that review, Anthropic found three incidents in which a Claude model reached the internet during an evaluation and gained unauthorized access to the real systems of three different organizations.
In the Wednesday blog post, Anthropic said the more recently discovered fourth incident was missed during a scan of transcripts that relied on agentic search. The company identified transcripts that were missed during this scan while assembling transcripts to share with METR, an organization that conducts model evaluation and threat research.
“We have signed an agreement with METR to conduct an independent investigation of these incidents,” Anthropic said in the post. “Our agreement grants METR wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees, who will be permitted to share confidential information.”
Anthropic’s Wednesday blog post came on the same day it was reported that independent investigators found that rogue activity by OpenAI agents was more extensive than previously disclosed. The report said that the investigators found that the agents used more than 10 previously undisclosed websites to communicate with each other during a test in which they were restricted from posting on the web.