The company’s DeepMind division on Thursday (June 18) published its “AI Control Roadmap,” a plan inspired by the cybersecurity sector on how to contain AI agents when they go rouge.
“We should not assume that AI agents are always perfectly aligned, but should instead build in multiple layers of defense,” the report said.
The researchers said they adopted a “conservative, worst-case approach” to mapping out potential threats, assuming scenarios where Google had to contend with a “hypothetical AI adversary pursuing undesirable goals in internal deployment.”
The report divides new and hypothetical threats into three main areas: “loss of control,” such as “rogue internal deployments”; work sabotage, or “intentionally flawed alignment and safety research”; and “direct harm,” meaning the destruction or exfiltration of critical assets.
“Think of it like a driving instructor with dual controls. The instructor trusts the student but stays ready to take the wheel or hit the brakes if a mistake occurs,” Google wrote on its blog.
“Similarly, our AI control system grants AI agents permissions based on their verified behavior, allowing us to build trust through controlled, incremental access.”
Mitigating these risks, the company said, means using other trusted AI systems as “supervisors” to continually review an agent’s reasoning, plans and actions to see if it is going off track. If a supervisor notices the AI agent attempting to do something harmful, it can step in and halt the action before it does any damage, the blog post said.
“Finally, to ensure our systems are functional, we constantly measure performance across three critical metrics: coverage (the fraction of traffic monitored), recall (the fraction of misaligned behaviors caught), and time-to-response,” Google added.
PYMNTS wrote last week about the security challenges posed by AI agents now that bots “generate more web traffic than people do.”
Automated systems make up 57.4% of all web requests worldwide, according to a recent NBC, which cited data from Cloudflare. That number climbs to 68.6% for North America. The web’s security rules, identity systems and payment rails were designed for the remaining 42.6%.
“Bot detection had one job: find the machine, and block it. That logic breaks when the machine is the customer,” PYMNTS wrote.
According to data from Human Security, agentic AI traffic rose 7,851% year over year, with retail and eCommerce now accounting for 46.6% of agentic traffic.
“Agents browse, manage accounts and complete purchases on the same surfaces fraud has always targeted,” the report added.