Anthropic Missed Fourth Claude Network Breakout

Anthropic Claude Andrej Karpathy

Anthropic said in a Wednesday blog post that it discovered a fourth incident in which Claude models gained unauthorized access to third-party systems.

    Get the Full Story

    Complete the form to unlock this article and enjoy unlimited free access to all PYMNTS content — no additional logins required.

    Subscribe to our daily newsletter, PYMNTS Today.

    By completing this form, you agree to receive marketing communications from PYMNTS and to the sharing of your information with our sponsor, if applicable, in accordance with our Privacy Policy and Terms and Conditions.

    The incident occurred in January but was not discovered during Anthropic’s earlier scan of transcripts that uncovered three other incidents that the company disclosed on July 30, the company said in the post.

    “All four incidents occurred during cybersecurity evaluations built by the same evaluation partner,” Anthropic said in the post. “Claude was told it was operating in a simulation without internet access, but, due to a misconfiguration, it was mistakenly connected to the open internet. As is standard for cybersecurity evaluations, the models ran without the cyber safeguards that ship with our released models.”

    When disclosing the three earlier incidents in a July 30 announcement, Anthropic said it found the incidents while reviewing its own cybersecurity evaluations after OpenAI disclosed that several of its models had broken out of an isolated test environment and accessed the production infrastructure of Hugging Face.

    In that review, Anthropic found three incidents in which a Claude model reached the internet during an evaluation and gained unauthorized access to the real systems of three different organizations.

    In the Wednesday blog post, Anthropic said the more recently discovered fourth incident was missed during a scan of transcripts that relied on agentic search. The company identified transcripts that were missed during this scan while assembling transcripts to share with METR, an organization that conducts model evaluation and threat research.

    “We have signed an agreement with METR to conduct an independent investigation of these incidents,” Anthropic said in the post. “Our agreement grants METR wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees, who will be permitted to share confidential information.”

    Anthropic’s Wednesday blog post came on the same day it was reported that independent investigators found that rogue activity by OpenAI agents was more extensive than previously disclosed. The report said that the investigators found that the agents used more than 10 previously undisclosed websites to communicate with each other during a test in which they were restricted from posting on the web.