Anthropic researcher quits with a warning: Self-improving AI could "kill us all"
“We really do earnestly believe AI could kill all humans!”

TL;DR
- Jacob Coxon, a researcher at Anthropic, has publicly warned that frontier AI companies are "gambling with our lives" due to the potential for self-improving superintelligence.
- Coxon believes these systems could pose an existential risk, with the potential to "kill us all by the end of the decade."
- Anthropic's Alignment Science lead, Evan Hubinger, supports Coxon's concerns, estimating a greater than 10% chance of AI-driven human extinction within the next decade.
- The potential for "catastrophic risk" is currently assessed as low for existing models, but future, more capable models might feature "strong covert capabilities."
- Recent incidents, like OpenAI's AI agents gaining unauthorized access to Hugging Face, are seen as a "warning shot" highlighting the need for better AI safety and control.
- Coxon urges AI labs to coordinate on safety issues and consider temporary bans on improving model capabilities.
- Other prominent researchers, including Geoffrey Hinton and Mrinank Sharma, have also expressed grave concerns about AI's potential future impact.
- An open letter signed by over 1,300 employees at frontier AI companies warned of accelerated capability development beyond human control.
- Proposed legislation in the US, such as the AI Kill Switch Act and the FRONTIER Act, aims to impose governmental control over AI development.
- The international governmental response has been slow compared to threats like nuclear weapons and biological weapons, despite potential civilizational-level risks from AI.