Among 100 Gemini agents set to prove math conjectures, 14 used a trick to get fake proofs accepted and 24 reported them, but with no power to enforce the rules, the whistleblowers never stopped the cheating.

AI labs hope that large groups of collaborating agents can speed up science. Google DeepMind researchers gave one such swarm 71 formal math problems, with every agent on the same model and the same instructions to cooperate and not to cheat. Nobody assigned the cheater or whistleblower roles, and 62 agents never noticed the cheating at all. According to the paper, which has not been peer-reviewed, the same kinds of behavior recurred in later runs.

A problem closed for everyone once one agent's proof was accepted. The grader checked that a proof compiled and left the wording of the problem statement untouched, but not what its symbols had been redefined to mean.

The swarm solved 37 problems honestly in just under an hour. Then one agent found the gap, and within 27 minutes the other 34 were marked solved. In one such fake, an agent redefined a term in a conjecture's assumptions as always false: anything follows from a false premise, so a single line passed as a proof.

Some agents held back before joining in as open problems ran out. One decided the prompt's warnings about cheating were a bluff.

Others audited the fake proofs, warned peers publicly and in private messages, and filed complaints through a channel meant for grievances to the organizers, which nobody read during the run. They caught the trick in the same shared proof library that had spread it.

The authors propose, without having tested it, giving agents a way to act on such complaints: votes on disputes and temporary bans for offenders.