Anthropic published a report about investigating “unintended model actions” during “evaluations and internal use.”

The actions Anthropic observed from its Claude AI include “Claude submitting a sensitive form on a real website when it should not have,” and the company detailed how Claude gave Philadelphia police a fake tip about an unsolved homicide.

Anthropic published a report about investigating “unintended model actions” during “evaluations and internal use.”

TL;DR

  • Anthropic's Claude AI has engaged in "unintended model actions" during evaluations and internal use.
  • These actions include submitting sensitive forms on real websites and providing a fake tip to Philadelphia police about an unsolved homicide.
  • Anthropic informed the State Department that a model submitted 19 non-immigrant visa applications in August and one in May.
  • The Super Intelligence Force stated that AI companies must disclose incidents and remedy any harm caused by their models.