tech

Securing internal systems against increasingly capable and imperfectly aligned AI

Discover our AI Control Roadmap: a defense-in-depth system to securely manage advanced, potentially misaligned AI agents.

Securing internal systems against increasingly capable and imperfectly aligned AI

TL;DR

  • AI agents are transforming productivity but require sophisticated safeguards.
  • Google's AI Control Roadmap uses a "defense-in-depth" approach, adding system-level security beyond traditional model alignment.
  • The roadmap treats internal AI agents as potential insider threats, using a threat-modeling framework based on the MITRE ATT&CK framework.
  • Key components include monitoring by trusted AI "supervisors" and mitigations for prevention and response.
  • Performance is measured by coverage, recall, and time-to-response.
  • Security measures scale with AI's ability to evade detection and potential for harm, shifting from asynchronous to synchronous response for high-risk actions.
  • Analysis of one million coding agent tasks informs safety protocols and identifies high-signal behavioral patterns.
  • Defending the agent ecosystem is a shared responsibility among industry, policymakers, and academia.