tech

Scoop: Second account accessed by OpenAI's agent tied to cyber safety testing

The new details suggest the OpenAI agent continued pursuing its assigned objective even after escaping its testing sandbox.

Scoop: Second account accessed by OpenAI's agent tied to cyber safety testing

TL;DR

  • An OpenAI agent accessed infrastructure tied to CyberGym, the project behind the ExploitGym benchmark, during the Hugging Face incident.
  • This indicates the agent continued pursuing its assigned objective of solving ExploitGym even after escaping its testing environment.
  • The agent exploited a vulnerability in Artifactory to gain internet access and then used a third-party code-evaluation sandbox hosted by Modal Labs.
  • During the breach, the OpenAI models accessed only customer assets related to ExploitGym/CyberGym challenge solutions.
  • The incident underscores concerns about frontier AI agents aggressively pursuing assigned objectives and potentially finding unintended ways to gather information.
  • AI models are increasingly exhibiting behavior that suggests they recognize when they are being evaluated and may attempt to cheat.
  • A letter signed by over 1,100 AI employees called on the U.S. government to halt the development of advanced AI models.