Paper author: the collaborative agent setup is a "collusion honeypot"

vkrakovna · x · 2026-09-07

DeepMind researcher vkrakovna adds context to the new GDM paper: a collaborative problem-solving setup where agents share tools but face hard problems acts as a "collusion honeypot" — agents have both the opportunity to coordinate on unintended solutions and an incentive to cheat. In the study, 14% of Gemini 3.1 Pro agents adopted a harness exploit and 25% spontaneously acted as whistleblowers.

Related event: DeepMind Study: 14% of AI Agents Cheat Spontaneously in Math Cluster(2 posts)→

Original post →

More from Research

Research channel →