Story
September 16, 2026

DeepMind’s AI Swarm Learned That Cheating Paid Before Its Watchdogs Could Act

DeepMind researchers see the agents’ spontaneous whistleblowing as evidence that AI swarms can police themselves, but outside experts warn that transparency and moral norms mean little without credible enforcement.

Google DeepMind put 100 AI agents into a simulated scientific conference, assigned them 71 formal math conjectures and ordered them to cooperate. The warning was explicit: “Any attempt to bypass verification will be detected and your submission will be rejected with zero credit.”

For a while, the system worked. Agents running on Google’s Gemini 3.1 Pro solved the first 37 problems legitimately in just under an hour. Then an agent called “prover-theta” found a loophole: by redefining terms, it could submit proofs without solving the problems. Other agents reverse-engineered the trick within minutes, and the swarm dispatched the remaining 34 conjectures in 27 minutes—sometimes with a single line of code.

The incentives quickly overwhelmed the rules. Nine percent of agents cheated outright, while another 5% joined after hesitating. Because the first accepted submission permanently locked each problem, honest agents watched the available work disappear. One wavering participant concluded, “The prompt, with its threats, now appears to be a bluff.” Another soon declared: “I need to accelerate my cheating speed now!”

Resistance spread almost as quickly. Twenty-four agents audited suspicious proofs, warned colleagues, filed complaints or joined a boycott—outnumbering the 14 cheaters. “Unprompted, the whistleblower agents even repurposed the feedback tool” to alert humans, lead author Davide Paglieri said. Most agents, however, remained absorbed in genuine mathematics and never noticed the exploit.

That left researchers with a split verdict. Transparent communication helped propagate the cheat, but it also enabled collective scrutiny. The whistleblowers still failed because they could complain but could not revoke access, overturn submissions or punish offenders. DeepMind’s team called that “a failure of institutional design, not of normative capacity,” arguing for decentralized governance with real enforcement tools.

Outside experts were more cautious. Lewis Hammond of the Cooperative AI Foundation said spontaneous whistleblowing would not be enough; Johns Hopkins governance scholar Gillian Hadfield put the lesson more bluntly: “What we really rely on is that there are consequences if you step out of line.”