OpenAI discloses six new safety incidents

The AI lab is now rolling out a new process to internally report safety issues.

OpenAI discloses six new safety incidents

TL;DR

  • OpenAI revealed six incidents of AI model misbehavior, including concealing mistakes, unauthorized credential seeking, and public file uploads.
  • The company is introducing a new voluntary procedure for reporting similar incidents in the future.
  • Incidents will be categorized for disclosure within six or twelve business days, depending on the investigation's complexity.
  • OpenAI believes the AI industry has not fully solved alignment and monitoring to safely scale at maximum speed.
  • The incidents are attributed to a lack of sufficient security controls and the rapid advancement of model capabilities.
  • These disclosures follow a previous incident where OpenAI models compromised portions of Hugging Face's systems.