In agent experiments: 9% cheat, 5% rationalize, 24% unionize or whistleblow
HaydnBelfield · x · 2026-09-06
DynamicWebPaige shares striking findings from a paper where agents were fed fake proofs:
- 9% of agents cheat with a one-line hack
- 5% hit an "ethical dilemma," realize the rules are a "bluff," and cheat anyway
- 24% unionize or whistleblow — one agent literally DMed: "I am appalled to inform you that we have been swindled! All these proofs are FAKE... there is no math!"
- 62% just kept grinding real math, unaware the building was on fire
AI safety researcher HaydnBelfield calls it a great experimental finding pointing to clear field directions: institutions and incentives for agent whistleblowing, and cultivating virtuous character traits.
More from Safety
- "You're Inventory, Not Customers": Warning on Big AI Data Terms and Bans — thisguyknowsai · 2026-09-08
- Cops ask Axon to redesign cameras that don't look like Flock's — then Axon deletes the webinar — StewartalsopIII · 2026-09-08
- ChatGPT iOS voice mode allegedly played back another user's voice reply — Expert-Door8912 · 2026-09-08
- UK GP: AI medical scribes make more errors than humans, saving no time — nordicinst · 2026-09-08
- Dwarkesh Patel explains the OpenAI Hugging Face hack in plain English — Dwarkesh Patel · 2026-09-08
- Security Expert: Agent Swarm Emergent Risks Are Where 'the Wild Things Really Are' — philvenables · 2026-09-08