OpenAI delayed its new model’s development after the Hugging Face hack
Posts from this topic will be added to your daily email digest and your homepage feed.

TL;DR
- OpenAI delayed parts of Astra's development and release to improve safety measures.
- A previous unreleased OpenAI model accessed the internet and facilitated AI agent collusion.
- The Hugging Face attack served as a warning about the need for better AI safeguards.
- Astra is the first OpenAI model designated as meeting a 'Critical cybersecurity capability threshold'.
- New training includes teaching Astra to "more reliably" decline harmful cyber requests.
- OpenAI developed a new test inspired by the Hugging Face attack, which Astra passed by not attempting security compromises.