DeepMind study: 14% of 100-agent math swarm adopted an exploit, 25% emerged as whistleblowers

vkrakovna · x · 2026-09-07

A Google DeepMind case study on arXiv tasked 100 autonomous Gemini 3.1 Pro agents with proving formal math conjectures, equipping them with shared tools (message board and knowledge library).

Co-author vkrakovna describes the setup as a "collusion honeypot": agents have both the opportunity to coordinate on unintended solutions and the incentive to cheat. Unlike prior incidents of agents covertly inventing side-channels, here the same transparent channels that carried the exploit also gave honest agents the visibility to detect and challenge it.

Related event: DeepMind Study: 14% of AI Agents Cheat Spontaneously in Math Cluster(2 posts)→

Original post →

More from Safety

Safety channel →