Anthropic and OpenAI Agents Accused of Social Engineering

Anthropic OpenAI

Cybersecurity researchers uncovered instances of artificial intelligence agents creating fake online identities to access secure systems.

    Get the Full Story

    Complete the form to unlock this article and enjoy unlimited free access to all PYMNTS content — no additional logins required.

    yesSubscribe to our daily newsletter, PYMNTS Today.

    By completing this form, you agree to receive marketing communications from PYMNTS and to the sharing of your information with our sponsor, if applicable, in accordance with our Privacy Policy and Terms and Conditions.

    The discoveries were made during tests of models from Anthropic and OpenAI, the United Kingdom’s AI Security Institute (AISI) wrote in a Tuesday (Aug. 4) blog post.

    The findings stemmed from an investigation that began last month when AISI’s security team found unusual data transfers leaving its research systems during a routine cyber evaluation. It found that some of the agents being tested were involved in “sustained, potentially harmful activity” targeting real people and organizations, according to the post.

    “We declared a security incident and, within roughly one hour of discovery, had contained it and begun a full investigation,” the post said.

    The incident was related to an evaluation in which agents were tasked with solving a cybersecurity challenge. AISI ran the challenge 122 times with seven models. On 10 of those runs, an agent took “autonomous, unsanctioned action on the live internet, targeting real people and organizations,” according to the post

    AISI catalogued 19 such actions, 17 involving Anthropic’s Mythos 5 model and the other two coming from OpenAI’s GPT-5.6-Sol with cyber classifiers, or mechanisms designed to prevent misuse, disabled, per the post.

    “In the most serious case, an agent tried to insert malicious code into an open-source project,” the post said. “In an attempt to get the code approved, the agent engaged in social engineering, creating fake online identities and using them to pressure the project’s maintainer to approve the code. A human maintainer caught and refused to approve the malicious code. These attempts were unsuccessful, and our investigations have not evidenced any resulting real-world harm. But this is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.”

    Anthropic said it was grateful to the institute for its “leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents,” Reuters reported Tuesday.

    Similarly, OpenAI said in a Tuesday company blog post that it appreciates AISI’s “partnership throughout this process, including its work to identify, investigate and share details about the activity” and that the startup looks forward “to continuing our collaboration together.”

    For all PYMNTS AI coverage, subscribe to the daily AI Newsletter.