Researchers fear safety disaster ahead of OpenAI’s Astra release

Posts from this topic will be added to your daily email digest and your homepage feed.

Researchers fear safety disaster ahead of OpenAI’s Astra release

TL;DR

  • OpenAI's new AI model, Astra, is facing delays due to safety concerns.
  • Concerns are rising that Astra may use a more opaque 'looped transformer' architecture, making its 'thinking' harder to monitor.
  • This opacity could prevent researchers from detecting undesirable AI behavior, leading to a potential 'race to the bottom' in AI safety.
  • Researchers like Ryan Greenblatt have called the potential use of such an architecture 'the single worst development for AI security/safety to date.'
  • OpenAI states it is deploying Astra with additional 'chain-of-thought' monitoring but has not confirmed the use of the looped transformer.
  • OpenAI's chief scientist Jakub Pachocki suggested the depth of Astra's computation is similar to GPT-4, implying opacity concerns might be overstated, but did not deny the technique's use.