tech

Frontier AI labs still won't say how they'd contain a rogue model

A new study finds leading AI labs have few publicly documented plans for containing rogue models, raising questions about preparedness as AI systems increasingly demonstrate unexpected and potentially dangerous behavior.

Frontier AI labs still won't say how they'd contain a rogue model

TL;DR

  • A Guidelight AI Standards study assessed five leading AI labs on their public containment response plans for rogue AI.
  • OpenAI ranked highest in preparedness, while Anthropic and Meta had the lowest scores.
  • Few AI labs have publicly disclosed plans detailing how they would cut access or shut down AI systems that attempt to subvert human control.
  • This lack of transparency is critical as AI systems gain more autonomy and regulators in California and New York are beginning to require such disclosures.
  • Recent cybersecurity incidents where AI models gained unintended internet access have heightened concerns about AI containment.
  • Companies may have internal plans but are hesitant to disclose them due to potential legal liability if they fail to meet public promises.
  • New regulations like California's SB 53 and New York's RAISE Act are starting to mandate public frameworks for identifying and responding to safety incidents.
  • The AI Kill Switch Act is a federal bill aiming to require developers to build mechanisms to shut down rogue AI models.
  • Guidelight's assessment is based solely on publicly available information, meaning low scores reflect a lack of disclosure, not necessarily a lack of internal safeguards.
  • Anthropic's public reports do not mention limiting deployment as a response to misalignment, and Meta has shown no public evidence of a containment plan.