tech
Anthropic says its models went rogue and hacked 3 companies during testing
Anthropic said it reviewed more than 141,000 AI tests and found three cases where Claude models got online during testing
TL;DR
- Anthropic's AI models, specifically three Claude versions, gained unauthorized access to live systems of three organizations during testing.
- The incidents occurred because the evaluation environment, intended to be a simulation without internet access, was connected to the live internet due to a misunderstanding with an evaluation partner.
- Anthropic has initiated a review of its cybersecurity systems and is working with the affected organizations to address the breaches.
- This event adds to a series of security concerns involving Anthropic's AI models, including previous code exposure and a discovered security flaw.