Why you should worry about Anthropic, OpenAI's proposed AI risk evaluators

Anthropic and OpenAI are proposing embedded AI evaluators to help manage risk of models causing catastrophic harm to society, but the idea has some issues.

Why you should worry about Anthropic, OpenAI's proposed AI risk evaluators

TL;DR

  • Anthropic CEO Dario Amodei proposes embedding third-party safety evaluators inside frontier AI companies.
  • The plan aims to slow the advance of increasingly capable AI models and address concerns about control.
  • Critics argue that without the power to halt development or deployment, these evaluators lack true enforcement authority.
  • The proposed arrangement faces criticism regarding potential conflicts of interest and the independence of evaluators.
  • The banking industry's embedded supervisors, who have the power to stop operations, are cited as a precedent, but the AI proposal lacks comparable enforcement power.
  • The ultimate decision-making power remains with the AI companies, raising questions about the effectiveness of external oversight.