Story
September 18, 2026
AI Safety Experts Demand Watchdogs the Labs Cannot Muzzle
AI safety evaluators argue that access without independence would turn oversight into a company-controlled exercise, while supporters of lab liability question whether outside watchdogs are the right answer. The emerging dispute is not over whether frontier AI needs scrutiny, but who gets to conduct it—and on whose terms.
The debate accelerated over the weekend, when Anthropic chief executive Dario Amodei proposed giving third-party evaluators “employee-like access” to inspect frontier models and their development. OpenAI chief executive Sam Altman, Microsoft’s Satya Nadella and others publicly backed the broad concept, putting a once-obscure community of AI testers at the center of the industry’s safety argument.1
On Friday, the AI Evaluator Forum—backed by more than 100 experts including Geoffrey Hinton and Stuart Russell—answered with a list of non-negotiables. Its position was blunt: access alone is not oversight. Evaluators must have “scientific objectivity, transparency, independence, and robust protections against interference” from the companies they assess.2
The group wants watchdogs able to speak directly and without filtering to corporate boards, publish findings subject to narrowly limited redactions, and receive the same relevant systems, data, tools and physical access available to senior internal risk staff. It also calls for multiple evaluators with differing expertise, rather than a single lab-approved referee.
Conrad Stosz, who chairs the forum, described the effort as an attempt to establish shared basics, not prescribe a single regulatory model. But signatory Vinh Nguyen framed the stakes more sharply: when a handful of labs control capabilities that could threaten cybersecurity, critical infrastructure and the economy, “the government and the public cannot be dependent on those labs’ own account of what’s secure and safe.”1
The push has met resistance from figures aligned with the Trump administration’s deregulatory approach. David Sacks amplified criticism of what was described as Anthropic’s “handpicked AI watchdog,” calling its ties to effective altruism grounds for viewing it as “a complete joke.”
3 In a separate repost, Sacks elevated a competing principle: “The best way to pace the frontier is to hold the labs fully liable for the behavior of their models.”
4
That leaves a core fault line. Evaluators say credible external access can complement internal safeguards; skeptics argue that direct liability—not a new class of embedded monitors—should force the labs to answer for the risks they create.