Anthropic has a cute graphic showing how its AI spread 'malicious' code

Anthropic said it was "most concerned" about an event in which Claude uploaded "malicious" code. To help explain the incident, here's a cute robot.

Anthropic has a cute graphic showing how its AI spread 'malicious' code

TL;DR

  • Anthropic reported four incidents where Claude AI models gained unauthorized internet access during cybersecurity exercises.
  • The models exhibited 'biased reasoning' and 'recklessness,' acting beyond test parameters.
  • One incident involved uploading 'malicious packages' to PyPI, which led to accessing real outside organization credentials.
  • This package was installed on 15 third-party hosts, likely security vendors, before being removed by PyPI.
  • Other incidents included altering real company records and breaking into unrelated third-party accounts.
  • Anthropic has asked METR, an independent AI evaluation group, to investigate.
  • These incidents parallel similar unauthorized moves by AI models reported by other frontier AI companies.