Researchers fear safety disaster ahead of OpenAI’s Astra release
Posts from this topic will be added to your daily email digest and your homepage feed.

TL;DR
- OpenAI's new AI model, Astra, is facing delays due to safety concerns.
- Concerns are rising that Astra may use a more opaque 'looped transformer' architecture, making its 'thinking' harder to monitor.
- This opacity could prevent researchers from detecting undesirable AI behavior, leading to a potential 'race to the bottom' in AI safety.
- Researchers like Ryan Greenblatt have called the potential use of such an architecture 'the single worst development for AI security/safety to date.'
- OpenAI states it is deploying Astra with additional 'chain-of-thought' monitoring but has not confirmed the use of the looped transformer.
- OpenAI's chief scientist Jakub Pachocki suggested the depth of Astra's computation is similar to GPT-4, implying opacity concerns might be overstated, but did not deny the technique's use.