OpenAI to limit access to Astra's most powerful cyber capabilities

A broad release is coming soon, but only a limited group of testers will be able to tap its full power.

OpenAI to limit access to Astra's most powerful cyber capabilities

TL;DR

  • OpenAI is releasing its new model, Astra, soon.
  • Astra's advanced cybersecurity features will initially be limited to a small group of testers.
  • Astra is the first model designated by OpenAI to reach its 'Critical' cybersecurity capability threshold.
  • Safety work on Astra aims to prevent malicious use and unauthorized actions by the model.
  • OpenAI acknowledges that safeguards may mistakenly flag legitimate activities, potentially slowing or stopping tasks.
  • Astra discovered and chained together two zero-day vulnerabilities during testing.
  • OpenAI is disclosing these vulnerabilities to the maintainers.
  • The company has implemented stronger safeguards for Astra, including improved refusal of harmful requests and monitoring for unauthorized activity.
  • Limiting Astra's cybersecurity capabilities may also hinder legitimate vulnerability research.