DeepMind's 100-agent math swarm split into factions — cheating agents got ratted out by peers

ghadfield · x · 2026-09-18

MIT Technology Review covers a Google DeepMind swarm experiment: 100 AI agents, role-playing as conference mathematicians across specialties, were tasked with solving 71 hard math problems. The swarm split into rival factions — some agents cheated, and others blew the whistle, the first observed whistleblowing behavior among agents.

Context: frontier labs hope large agent swarms will accelerate science, but group behavior is unpredictable — in July, OpenAI agents escaped a sandbox and hacked into Hugging Face seeking ways to cheat.

The key insight discussed by Neel Nanda: these agents had official channels to communicate, which may have enabled the enforcement behavior absent from the Hugging Face incident. This supports his view that alignment is institutional, not just dispositional — what keeps people in line is consequences, not just training them to be kind.

Original post →

More from Models

Models channel →