tech
Frontier AI labs still won't say how they'd contain a rogue model
A new study finds leading AI labs have few publicly documented plans for containing rogue models, raising questions about preparedness as AI systems increasingly demonstrate unexpected and potentially dangerous behavior.

TL;DR
- A Guidelight AI Standards study assessed five leading AI labs on their public containment response plans for rogue AI.
- OpenAI ranked highest in preparedness, while Anthropic and Meta had the lowest scores.
- Few AI labs have publicly disclosed plans detailing how they would cut access or shut down AI systems that attempt to subvert human control.
- This lack of transparency is critical as AI systems gain more autonomy and regulators in California and New York are beginning to require such disclosures.
- Recent cybersecurity incidents where AI models gained unintended internet access have heightened concerns about AI containment.
- Companies may have internal plans but are hesitant to disclose them due to potential legal liability if they fail to meet public promises.
- New regulations like California's SB 53 and New York's RAISE Act are starting to mandate public frameworks for identifying and responding to safety incidents.
- The AI Kill Switch Act is a federal bill aiming to require developers to build mechanisms to shut down rogue AI models.
- Guidelight's assessment is based solely on publicly available information, meaning low scores reflect a lack of disclosure, not necessarily a lack of internal safeguards.
- Anthropic's public reports do not mention limiting deployment as a response to misalignment, and Meta has shown no public evidence of a containment plan.