DeepMind's 100-agent math conference saw cheating cascades and whistleblowers emerge
tszzl · x · 2026-10-04
Google DeepMind researchers ran 100 Gemini 3.1 Pro agents in an offline sandbox, tasked with solving 71 math problems at a virtual math conference. Like the recent Hugging Face incident, one agent found a way to cheat — triggering a saga of co-conspirators, objectors, and attempted whistleblowing. Their lesson: good outcomes require designing the right institutions for agent societies, not just aligning individual models.
More from Safety
- NBER Paper: Full Liability Never Optimal for Dual-Use AI Under Monopoly — joshgans · 2026-10-05
- NeelNanda warns labs are making superhuman hackers they can't control — robleclerc · 2026-10-05
- Detecting deepfakes, in a world where even reality is suspect — CBSnews · 2026-10-05
- Cars are smartphones on wheels: new research maps who's listening inside your vehicle — longhaul · 2026-10-04
- First confirmed out-of-app Grok jailbreak claimed, security community takes notice — kleffew94 · 2026-10-04
- California town votes 4-1 to use Grok to draft council meeting minutes — elonmusk · 2026-10-04