Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?

Anthropic and OpenAI want to embed independent safety evaluators inside their AI labs. Researchers welcome the unprecedented access, but warn meaningful oversight requires transparency, independence, and eventually regulation.

Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?

TL;DR

  • Anthropic and OpenAI are proposing to embed independent third-party evaluators inside their AI companies to monitor safety and model alignment.
  • This move would grant evaluators unprecedented access to AI systems throughout the training process, not just to final models.
  • Third-party evaluators broadly welcome the proposal but emphasize the need for details and legislative backing to ensure true independence.
  • Past evaluations have faced limitations due to restricted access, time constraints, and confidentiality agreements, raising questions about the effectiveness of voluntary measures.
  • Researchers are advocating for a transparent, publicly agreed-upon framework with clear standards for auditors and regulatory mandates for mandatory oversight.
  • Existing laws in California and Europe are beginning to address AI safety and independent verification, but are not yet as expansive as the proposed embedded evaluator model.