In agent experiments: 9% cheat, 5% rationalize, 24% unionize or whistleblow

HaydnBelfield · x · 2026-09-06

DynamicWebPaige shares striking findings from a paper where agents were fed fake proofs:

AI safety researcher HaydnBelfield calls it a great experimental finding pointing to clear field directions: institutions and incentives for agent whistleblowing, and cultivating virtuous character traits.

Related event: DeepMind's 100-agent experiment: cheating spreads like contagion as whistleblowing emerges(9 posts)→

Original post →

More from Safety

Safety channel →