Anthropic and OpenAI need truly independent safety evaluators, experts say in public letter

A coalition of over 100 AI experts are urging independence and transparency from Anthropic, OpenAI and other foundation model labs to conduct evaluations.

Anthropic and OpenAI need truly independent safety evaluators, experts say in public letter

TL;DR

  • Over 100 AI experts and evaluators are demanding necessary resources and protections for independent AI safety testing.
  • The group wants to ensure AI companies provide scientific objectivity, transparency, independence, and robust protections for third-party evaluators.
  • This initiative aims to hold foundation model providers accountable for their pledges to support third-party AI safety testing.
  • Key proposals include independent ownership, avoiding commercial conflicts, and accepting payment only for work performed, not for specific findings.
  • Evaluators should have access equivalent to internal employees, including candid communication and sensitive internal data, with limited exceptions.
  • Companies should facilitate transparency, limit NDAs, and allow prompt communication with boards and public release of findings.
  • Evaluators need protection from retaliation, including from retaliatory litigation.
  • The efforts complement, not replace, internal evaluation efforts and broader external oversight by companies.