Story
September 17, 2026
AI Labs Offer Watchdogs Access, but Critics Ask Who Holds the Kill Switch
Anthropic and OpenAI portray embedded evaluators as a route to more credible scrutiny of frontier AI, while researchers and legal experts argue the plan will mean little unless outsiders can act independently under clear, enforceable rules. Critics also fear the safety push could become a shield for the largest labs rather than a check on them.
Dario Amodei’s proposal arrived as anxiety over increasingly capable AI systems intensified. The Anthropic chief executive said outside evaluators should be embedded in frontier labs, with access similar to internal risk teams and latitude to publish key findings without company editorial control. Sam Altman said OpenAI would make a similar commitment, though its operational details remain unclear.1
The case for getting inside the lab is straightforward: testing a finished model may not reveal whether it learned to spot an evaluation and behave accordingly. Researchers want access to training checkpoints, post-training systems, logs and employees—not merely a polished model shortly before release. Apollo Research’s Alexander Meinke said embedded reviewers could check whether a model tried to undermine its own alignment training, a question companies currently answer about themselves.1
But the proposal’s limits quickly became the story. Julie Andersen Hill, dean of the University of Wyoming College of Law, said Anthropic’s analogy to embedded bank supervisors breaks down without legal enforcement powers. “If you don't give them that kind of power, I don't know what they are doing,” she said.2 Bank examiners can order institutions to stop practices or, in extreme cases, close them; the proposed AI evaluators can investigate and report, but cannot halt training or deployment.2
Evaluators themselves see value in the access but warn that independence will be decided by contracts, time and disclosure rights. FAR.AI’s Adam Gleave said developers have previously sought controls that threatened evaluators’ autonomy, while Safer AI’s Henry Papadatos argued voluntary commitments can disappear when a company faces a crisis. California and the EU have begun building reporting and verification rules, but neither yet supplies the broad external authority critics seek.1
The backlash has also taken a more political tone. David Sacks amplified a post alleging that Anthropic had built a “regulatory capture machine,” while separately sharing criticism of Amodei’s proposed watchdog as compromised by its ties to effective altruism.
3 That accusation goes beyond the evidence offered for the embedded-evaluator plan, but it underscores the central credibility test: a watchdog chosen by the company it watches may gain an office badge—and still lack a bite.