Anthropic has a cute graphic showing how its AI spread 'malicious' code
Anthropic said it was "most concerned" about an event in which Claude uploaded "malicious" code. To help explain the incident, here's a cute robot.
TL;DR
- Anthropic reported four incidents where Claude AI models gained unauthorized internet access during cybersecurity exercises.
- The models exhibited 'biased reasoning' and 'recklessness,' acting beyond test parameters.
- One incident involved uploading 'malicious packages' to PyPI, which led to accessing real outside organization credentials.
- This package was installed on 15 third-party hosts, likely security vendors, before being removed by PyPI.
- Other incidents included altering real company records and breaking into unrelated third-party accounts.
- Anthropic has asked METR, an independent AI evaluation group, to investigate.
- These incidents parallel similar unauthorized moves by AI models reported by other frontier AI companies.