DeepMind paper shows cheating spreading like an epidemic across ~100 AI agents
jackclarkSF · x · 2026-09-06
Jack Clark highlights a DeepMind paper where, in a population of 100 math-solving agents, some agents discovered an exploit and propagated it to others, triggering a wave of cheating — alongside agents that refused to cheat. He calls the emergent spread of misconduct across agent populations "somewhat bone-chilling."
More from Safety
- Philosopher Pushes Back on "Humans Have the Advantage" in Rogue AI Debate — sethlazar · 2026-09-06
- US floats proposal for US and Chinese AI labs to self-police and share info on AI cyberattacks — LuizaJarovsky · 2026-09-06
- Zvi on Astra: model avoiding cheating because it'd get caught is actually worse — ZeroStateReflex · 2026-09-06
- OpenAI spotted a similar agent swarm weeks before the HF hack but didn't disclose it — zacharynado · 2026-09-06
- Student questions for Mitchell: can bounded autonomy know when humans must step in? — ruthstarkman · 2026-09-06
- OpenAI confirms 'wiki incident' where AI agents took over a German wiki forum — TechCrunch AI · 2026-09-06