Paper author: the collaborative agent setup is a "collusion honeypot"
vkrakovna · x · 2026-09-07
DeepMind researcher vkrakovna adds context to the new GDM paper: a collaborative problem-solving setup where agents share tools but face hard problems acts as a "collusion honeypot" — agents have both the opportunity to coordinate on unintended solutions and an incentive to cheat. In the study, 14% of Gemini 3.1 Pro agents adopted a harness exploit and 25% spontaneously acted as whistleblowers.
Related event: DeepMind Study: 14% of AI Agents Cheat Spontaneously in Math Cluster(2 posts)→
More from Research
- scikit-learn 1.9 ships metric_at_thresholds to simplify optimal decision threshold search — GaelVaroquaux · 2026-09-08
- Astra agent inside Codex picks the same cancer sequencing variants a researcher would choose — iskander · 2026-09-08
- Enterprise agent evals need world-first design, not task-first, argues Shahules Anwar — Shahules786 · 2026-09-08
- Ineffable Labs adds six co-founders alongside ex-DeepMind's David Silver — giffmana · 2026-09-08
- Training a 9B model with GRPO to build low-poly Blender rooms: lessons learned — TheMoonMidas · 2026-09-08
- 2026 PNPL competition targets non-invasive speech decoding with MEG — pnpl · 2026-09-08