Story
September 19, 2026

AI Safety Watchdogs Demand Freedom From the Labs They Inspect

AI safety evaluators argue that access to frontier labs means little without independence and protection from retaliation, while their critics see the emerging watchdog system as vulnerable to ideological capture. The dispute is not over whether powerful models need scrutiny, but over who can credibly deliver it.

The debate accelerated after Anthropic chief executive Dario Amodei proposed giving outside safety evaluators employee-like access to frontier-model development, an approach OpenAI chief executive Sam Altman and other industry figures have backed. For evaluators, however, access is only the opening bid: they say a reviewer chosen, paid or constrained by the company under review cannot serve as a meaningful check.

On Friday, the AI Evaluator Forum and more than 100 signatories, including Geoffrey Hinton and Stuart Russell, set out their minimum conditions. Their letter calls for evaluators to have “scientific objectivity, transparency, independence, and robust protections” from interference, alongside access to sensitive systems, data, tools and staff comparable to that available to internal assessors. Conrad Stosz, the forum’s chair, said such access could let reviewers inspect unreleased systems and internal data, giving the public greater confidence about risks the labs may not disclose on their own.

The coalition’s model is deliberately more demanding than a friendly audit. It wants multiple evaluators with differing expertise, direct communication with boards, limited nondisclosure constraints and the right to publish findings after narrowly tailored, time-limited redactions. It also insists that watchdogs be protected against retaliatory lawsuits or funding cuts when their conclusions embarrass a host company. Microsoft chief executive Satya Nadella has welcomed the embedded-evaluator concept, according to reporting on the letter.

But the plan has immediately exposed a second accountability problem: who watches the watchdogs? David Sacks amplified a New York Post attack describing Anthropic’s purported watchdog as “handpicked” and tied to the Effective Altruism movement, calling the arrangement “a complete joke.” In a separate repost, he elevated Naval Ravikant’s alternative: “The best way to pace the frontier is to hold the labs fully liable for the behavior of their models.”

That leaves the central fault line stark. The evaluators do not present themselves as a replacement for internal safety work or broader oversight; they want a protected window into the labs. Their critics fear that window could simply install another unaccountable power center.