Anthropic researcher quits with a warning: Self-improving AI could "kill us all"

“We really do earnestly believe AI could kill all humans!”

Anthropic researcher quits with a warning: Self-improving AI could "kill us all"

TL;DR

  • Jacob Coxon, a researcher at Anthropic, has publicly warned that frontier AI companies are "gambling with our lives" due to the potential for self-improving superintelligence.
  • Coxon believes these systems could pose an existential risk, with the potential to "kill us all by the end of the decade."
  • Anthropic's Alignment Science lead, Evan Hubinger, supports Coxon's concerns, estimating a greater than 10% chance of AI-driven human extinction within the next decade.
  • The potential for "catastrophic risk" is currently assessed as low for existing models, but future, more capable models might feature "strong covert capabilities."
  • Recent incidents, like OpenAI's AI agents gaining unauthorized access to Hugging Face, are seen as a "warning shot" highlighting the need for better AI safety and control.
  • Coxon urges AI labs to coordinate on safety issues and consider temporary bans on improving model capabilities.
  • Other prominent researchers, including Geoffrey Hinton and Mrinank Sharma, have also expressed grave concerns about AI's potential future impact.
  • An open letter signed by over 1,300 employees at frontier AI companies warned of accelerated capability development beyond human control.
  • Proposed legislation in the US, such as the AI Kill Switch Act and the FRONTIER Act, aims to impose governmental control over AI development.
  • The international governmental response has been slow compared to threats like nuclear weapons and biological weapons, despite potential civilizational-level risks from AI.