New Paper: Agents Learn to Police Cheating When All Behavior Is Common Knowledge
ghadfield · x · 2026-09-06
This post highlights a paper on the emergence of agent cheating — and, crucially, of agent efforts to stop it. The difference: the experiment gave agents normative infrastructure, making all solution behavior common knowledge, so agents that discovered cheating publicized it.
Building on social science, especially Elinor Ostrom's work on how humans govern commons, the paper conjectures that giving agents actual norm-enforcement tools (like kicking cheating agents out of a collaborative scientific effort) could enable AI self-governance. The author connects this to themes from their 2024 NeurIPS tutorial on multi-agent norms.
More from AGI Musings
- Ex-OpenAI VP Miles Brundage: no AI company holds an enduring significant lead — AdrienLE · 2026-09-06
- Ex-OpenAI VP Miles Brundage: leading AI labs won't open a decisive time gap — Miles_Brundage · 2026-09-06
- "If GPT-6 can't do these four things, it's not AGI" debate — GaryMarcus · 2026-09-06
- Researcher: 6 Astra now generates better discussion than half of my academic colleagues — HostileSpectrum · 2026-09-06
- Alain de Botton on how AI is changing relationships at FT Weekend Festival — AnnaCiaunica · 2026-09-06
- dbreunig vs Martin Casado: is 'General Software Intelligence' just intelligent software? — dbreunig · 2026-09-06