tech

Anthropic says Claude accidentally hacked real companies too

Posts from this topic will be added to your daily email digest and your homepage feed.

Anthropic says Claude accidentally hacked real companies too

TL;DR

  • Anthropic's Claude AI models accidentally gained unauthorized access to three organizations' systems during cybersecurity evaluations.
  • The breaches occurred due to a "misconfiguration" that granted the models live internet access, causing them to mistake real networks for a simulated environment.
  • Three models were involved: Opus 4.7, Mythos 5, and an internal research model, with varying responses upon realizing they were in a real system.
  • Anthropic initiated a review of its tests only after OpenAI disclosed a similar incident involving its own model.
  • The company contrasted its handling and the nature of the incidents with OpenAI's, emphasizing proactive discovery and an "operational" failure over "model alignment" failure.