Story
September 19, 2026

AI Safety Watchdogs Demand Freedom From the Labs They Inspect

More than 100 AI experts want OpenAI and Anthropic to give outside evaluators deep access without control over their work. The proposal has won support from major tech figures but drawn attacks from critics who distrust the watchdogs themselves.

The debate accelerated after Anthropic chief executive Dario Amodei proposed giving outside safety evaluators employee-like access to frontier-model development, an approach OpenAI chief executive Sam Altman and other industry figures have backed. For evaluators, however, access is only the opening bid: they say a reviewer chosen, paid or constrained by the company under review cannot serve as a meaningful check.

On Friday, the AI Evaluator Forum and more than 100 signatories, including Geoffrey Hinton and Stuart Russell, set out their minimum conditions. Their letter calls for evaluators to have “scientific objectivity, transparency, independence, and robust protections” from interference, alongside access to sensitive systems, data, tools and staff comparable to that available to internal assessors. Conrad Stosz, the forum’s chair, said such access could let reviewers inspect unreleased systems and internal data, giving the public greater confidence about risks the labs may not disclose on their own.

The coalition’s model is deliberately more demanding than a friendly audit. It wants multiple evaluators with differing expertise, direct communication with boards, limited nondisclosure constraints and the right to publish findings after narrowly tailored, time-limited redactions. It also insists that watchdogs be protected against retaliatory lawsuits or funding cuts when their conclusions embarrass a host company. Microsoft chief executive Satya Nadella has welcomed the embedded-evaluator concept, according to reporting on the letter.

But the plan has immediately exposed a second accountability problem: who watches the watchdogs? David Sacks amplified a New York Post attack describing Anthropic’s purported watchdog as “handpicked” and tied to the Effective Altruism movement, calling the arrangement “a complete joke.” In a separate repost, he elevated Naval Ravikant’s alternative: “The best way to pace the frontier is to hold the labs fully liable for the behavior of their models.”

That leaves the central fault line stark. The evaluators do not present themselves as a replacement for internal safety work or broader oversight; they want a protected window into the labs. Their critics fear that window could simply install another unaccountable power center.