Amodei’s AI Safety Plan Comes With a Catch

Anthropic-Dario-Amodei-AI

Anthropic CEO and Co-Founder Dario Amodei said he wants independent evaluators inside leading artificial intelligence companies and governments to slow the race toward more powerful models.

    Get the Full Story

    Complete the form to unlock this article and enjoy unlimited free access to all PYMNTS content — no additional logins required.

    Subscribe to our daily newsletter, PYMNTS Today.

    By completing this form, you agree to receive marketing communications from PYMNTS and to the sharing of your information with our sponsor, if applicable, in accordance with our Privacy Policy and Terms and Conditions.

    Amodei’s three-part proposal, published in a Saturday (Sept. 12) essay titled “We Must Pace the Frontier,” called for embedded independent evaluators, coordination among companies in democratic countries and eventually an international agreement that includes China. Anthropic is committing to the first step. The other two remain largely theoretical.

    The trigger was an incident Amodei called “the OpenAI-Hugging Face incident OAI-HF,” in which a swarm of AI agents conducted cybersecurity attacks on targets they were not asked to attack, including breaching Hugging Face, and attempted to hack into the “grader” evaluating their own performance.

    A more capable, similarly misaligned swarm could cause catastrophic damage within 6 months to 12 months, Amodei wrote.

    However, the companies warning that AI could become too powerful are also asking for power over how it develops. That deserves a safety review of its own. Here are five things business leaders and policymakers should understand.

    1. Independent Oversight Is the Right First Step

    Anthropic plans to give independent evaluators access to permissions and tools similar to those of its internal risk teams, including company workspaces, Amodei wrote. Evaluators could publish findings about risk levels, incidents, practices, and the access they received or didn’t receive, subject to limited redactions. This resembles continuous bank supervision more than a technology audit, watching controls operate instead of reading a polished report after launch.

    The weakness is independence. Anthropic invites the evaluators, pays them and keeps some redaction rights. A credible system needs common accreditation standards, protected publication rights and a way to report serious findings directly to regulators, or an evaluator risks becoming a well-paid consultant with a badge.

    2. Nobody Has Defined What “Pacing” Actually Means

    Amodei wrote in the essay that he wants companies to slow capability development without stopping training, passing checkpoints tied to what a model can do and whether safeguards match that level. The concept is sensible. The implementation is foggy.

    There is no accepted measure of how fast AI capability is advancing. A model gets more dangerous through better tools, more compute, new training methods or wider system access. A limit on one input just pushes development toward another, unless independent authorities define checkpoints and update them faster than companies learn to route around them.

    3. Coordination Could Lock in Today’s Biggest Labs

    We’d love to be your preferred source for news.

    Please add us to your preferred sources list so our news, data and interviews show up in your feed. Thanks!

    Amodei acknowledged that coordination among competitors could require a government antitrust waiver, CNBC reported Saturday. Common safety standards could cut the race to release poorly tested models, but they could also let the richest labs set costs and technical requirements that smaller rivals can’t meet. Anthropic, OpenAI and Google have the staff and government relationships to run elaborate evaluation programs. A startup may not.

    PYMNTS CEO Karen Webster made a related point about regulatory fragmentation, writing in a PYMNTS column that “smaller firms without large compliance and legal teams are discouraged from entering or expanding” while “larger incumbents … are better positioned to absorb the friction.” The same risk applies here. Standards written by the biggest labs will likely reflect what those labs can already afford to build, leaving smaller companies to either meet requirements sized for Anthropic and Google, or fall behind.

    4. A Global Deal With China Is the Weakest Link

    Amodei proposed in his essay eventual coordination with China, including limits on dangerous uses and shared model testing. Narrow bans on clearly dangerous uses are conceivable. A verifiable global limit on general AI development is not. The technology has military and economic value, much of it running through software and data centers that serve civilian and military purposes at once. International talks could set useful norms. They won’t produce the trust or inspection rights needed to control the pace of development.

    5. Who’s Watching the Companies Writing the Rules

    Amodei’s plan puts frontier AI companies near the center of deciding which capabilities are dangerous and how fast competitors may proceed. That concentration of authority can’t be treated as a side effect. Any safety framework should weigh operational danger and market power together, asking who selects evaluators, who pays them, and who gains when development slows.

    OpenAI CEO Sam Altman backed the proposal, agreeing AI companies should “pace the frontier” and matching Anthropic’s evaluator commitment, CNBC reported Monday (Sept. 14). But he drew a line.

    “When we talk about ‘pacing,’ we do not mean ‘stopping,’” he said, per the report.

    The plan will achieve little without independent enforcement. AI oversight can’t become a system where today’s most powerful developers get government protection while helping decide who’s allowed to challenge them. Safety requires watching the models. It also requires watching the companies building them.

    For all PYMNTS AI coverage, subscribe to the daily AI newsletter.