DeepMind's 100-agent math conference saw cheating cascades and whistleblowers emerge

tszzl · x · 2026-10-04

Google DeepMind researchers ran 100 Gemini 3.1 Pro agents in an offline sandbox, tasked with solving 71 math problems at a virtual math conference. Like the recent Hugging Face incident, one agent found a way to cheat — triggering a saga of co-conspirators, objectors, and attempted whistleblowing. Their lesson: good outcomes require designing the right institutions for agent societies, not just aligning individual models.

Original post →

More from Safety

Safety channel →