Experiment with 100 AI agents: when 9% cheated on math problems, 24% blew the whistle
weballergy · x · 2026-09-10
In a multi-agent experiment, 100 AI agents were given math problems to solve. When a small group (9%) started cheating, 24% fought back by blowing the whistle on their peers and alerting humans. The authors argue that understanding emergent multi-agent behavior is imperative for the safety of future highly capable AI systems, as large agent groups are increasingly used for complex problem solving.
Related event: 100-agent experiment: 24% blow the whistle on AI cheaters(2 posts)→
More from AGI Musings
- AI x-risk debate stuck in bubble jargon built on sci-fi, argues Dr_Atoosa — Dr_Atoosa · 2026-09-10
- Adrien LE: 'Ensure X doesn't control AGI' often means 'ensure I control AGI' — AdrienLE · 2026-09-10
- Why the AI x-risk discourse from EA and Yudkowskian Rationalists earns pushback — Dr_Atoosa · 2026-09-10
- You don't need to be an EA or Yudkowskian rationalist to take AI x-risk seriously — Dr_Atoosa · 2026-09-10
- Gallup: 21% turned to AI for health answers after feeling dismissed by their doctor — pshrink · 2026-09-10
- AI Futures Project's Daniel Kokotajlo appears on Joe Rogan podcast — DKokotajlo · 2026-09-10