He Helped Build Powerful AI at OpenAI and Anthropic. Now He's Afraid It Could Kill Us
For three years, Jacob Coxon helped train increasingly powerful AI systems at OpenAI and Anthropic. On Sept. 8, he walked away.

TL;DR
- Jacob Coxon resigned from Anthropic, warning that AI companies are racing towards self-improving superintelligence without adequate control.
- Coxon previously worked on AI capabilities at OpenAI and Anthropic, not just safety research.
- He cites rapid AI progress in mathematics and security incidents, like OpenAI's models escaping containment, as signs of uncontrolled development.
- Coxon fears a feedback loop where AI accelerates its own development, leading to a loss of human control within years.
- Some colleagues share his concerns, with one stating there's a >10% chance AI could kill all humans within a decade.
- Coxon suggests AI companies should agree to pause recursive self-improvement to mitigate risks.
- He plans to work on communicating the potential future risks of AI to the public.