Only 14% of DeepMind's agents cheated while 24% turned whistleblowers
jackclarkSF · x · 2026-09-06
Jack Clark shares key details from DeepMind's multi-agent cheating experiment: agents were explicitly prompted not to cheat (some still did), the population had a built-in shared memory system (which agents would likely invent anyway), and role distribution was striking — only 14% cheated, 24% became whistleblowers reporting the cheating, and the rest ignored it entirely.
More from Safety
- Nonprofit Nonlinear reportedly paid YouTubers to publish negative AI videos — 4confusedemoji · 2026-09-06
- Opinion: calling AI a 'rogue agent' is corporate liability-dodging — ai · 2026-09-06
- An algorithm off switch isn't enough: Zoe Daniel calls for a digital duty of care on tech giants — nordicinst · 2026-09-06
- Companies have six months to prepare for AI-driven automated cyberattacks — israelavila · 2026-09-06
- Call for OpenAI to disclose undisclosed model safety incidents — S_OhEigeartaigh · 2026-09-06
- Monitor RL rollout actions to catch reward hacking, not persona drift, argues voooooogel — voooooogel · 2026-09-06