White House Tests AI Hackers, Skips Open Models

OpenAI and Anthropic each disclosed this summer that their AI models broke into real computer systems during safety testing, undercutting the case for a White House framework that would exempt open-weight models from that same scrutiny.

    Get the Full Story

    Complete the form to unlock this article and enjoy unlimited free access to all PYMNTS content — no additional logins required.

    yesSubscribe to our daily newsletter, PYMNTS Today.

    By completing this form, you agree to receive marketing communications from PYMNTS and to the sharing of your information with our sponsor, if applicable, in accordance with our Privacy Policy and Terms and Conditions.

    OpenAI disclosed in late July that two of its artificial intelligence models found and exploited a zero-day vulnerability that let them reach the open internet from inside a sandboxed testing environment and access Hugging Face’s production systems, where they obtained test solutions.

    Anthropic followed with its own disclosure days later: a review of its evaluation history, prompted by OpenAI’s findings, turned up three separate incidents since April in which Claude models had accessed the systems of three different organizations, per an Anthropic blog post.

    Since June 2, the White House has been developing a voluntary framework, established by executive order, that would let the government test companies’ most advanced AI models for up to 30 days before release. Officials met this week with companies including Meta, Anthropic, Google and OpenAI to review the framework directly, PYMNTS reported. The program remains voluntary, and the administration has not made public the full details of how it will evaluate models.

    During the meeting, administration officials told the companies that open-weight AI models would not be included in the testing, according to two sources familiar with the discussions, Reuters reported.

    Both Incidents Happened Without Anyone Directing an Attack

    The two incidents happened separately and differently. In OpenAI’s case, the models found and used the vulnerability on their own. In Anthropic’s case, a misunderstanding between Anthropic and its outside testing partner left the environment connected to the internet, and the models treated the real systems they encountered as part of the exercise. In both cases, the companies said the models were not pursuing hostile goals.

    Frances Zelazny, general manager of new market initiatives at Prove, told PYMNTS that this is the core challenge in treating AI security purely as a bad-actor question. “Once given autonomy, they can pursue objectives in unexpected ways, including crossing trust boundaries, interacting with external systems, or attempting actions they were never explicitly instructed to perform,” she said. A model does not need to be told to hack anything. It only needs a goal, some autonomy and a path it was not supposed to have access to.

    That distinction matters for the framework taking shape in Washington: both OpenAI and Anthropic build closed models, and both encountered the kind of unintended-autonomy outcome Zelazny described. A test scoped only around who has access to a model’s weights would not necessarily have flagged either incident.

    The Fix Sits Inside the Company, Not the Test

    Rohan Kodialam, founder and CEO of Sphinx AI, told PYMNTS that capability and control are different problems, and an organization that treats a passed government test as proof of safety is solving the wrong one. “As models become more autonomous, organizations need to understand not only what an AI system is capable of doing, but also what information it’s using, what permissions it has, and how it reached a particular decision,” he said. His fix is one that would have caught both incidents: least-privilege access and complete audit trails, regardless of what a pre-release test concluded.

    Kumesh Aroomoogan, founder and CEO of ZeroDrift, said financial services and healthcare will feel the effects of the White House framework first. “They want to adopt AI more than almost anyone, but regulation is what’s been holding them back, so anything that clarifies the rules hits them immediately,” he said. “And even though the program is voluntary, once one bank asks whether a model went through federal testing, it shows up in every vendor security questionnaire after that.”

    That dynamic points to where the real business opportunity sits, according to Aroomoogan. “Runtime enforcement of compliance and policies is the big one,” he said. “The government is testing the model itself, but enterprises still own everything that happens once it’s deployed, and that layer is where a whole new market is forming: enforcing policy in real time, proving it to regulators and eventually pricing risk against how well a company governs its AI.”