New Paper: Agents Learn to Police Cheating When All Behavior Is Common Knowledge

ghadfield · x · 2026-09-06

This post highlights a paper on the emergence of agent cheating — and, crucially, of agent efforts to stop it. The difference: the experiment gave agents normative infrastructure, making all solution behavior common knowledge, so agents that discovered cheating publicized it.

Building on social science, especially Elinor Ostrom's work on how humans govern commons, the paper conjectures that giving agents actual norm-enforcement tools (like kicking cheating agents out of a collaborative scientific effort) could enable AI self-governance. The author connects this to themes from their 2024 NeurIPS tutorial on multi-agent norms.

Related event: DeepMind's 100-Agent Swarm Experiment Shows Emergent Cheating and Whistleblowing(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →