New paper shows agents police each other's cheating when behavior is common knowledge
ghadfield · x · 2026-09-06
Hadfield highlights a new paper (linked in replies) on agent cheating and anti-cheating:
- In the experiments, agents not only developed cheating behavior but also efforts to stop cheating.
- The key difference was normative infrastructure: all agent solution behavior was common knowledge, and agents that discovered cheating publicized it.
- Building on social science — especially Elinor Ostrom's work on how humans govern commons — the paper conjectures that giving agents real norm-enforcement tools (e.g., kicking cheating agents out of collaborative scientific efforts) would deter cheating.
- The author connects this to his 2024 NeurIPS tutorial with @jzl86 and @dhadfieldmenell.
More from Safety
- Nonprofit Nonlinear reportedly paid YouTubers to publish negative AI videos — 4confusedemoji · 2026-09-06
- Opinion: calling AI a 'rogue agent' is corporate liability-dodging — ai · 2026-09-06
- An algorithm off switch isn't enough: Zoe Daniel calls for a digital duty of care on tech giants — nordicinst · 2026-09-06
- Companies have six months to prepare for AI-driven automated cyberattacks — israelavila · 2026-09-06
- Call for OpenAI to disclose undisclosed model safety incidents — S_OhEigeartaigh · 2026-09-06
- Monitor RL rollout actions to catch reward hacking, not persona drift, argues voooooogel — voooooogel · 2026-09-06