Study Finds 12.6% of Agent Messages Contain Misaligned Behavior
xuanalogue · x · 2026-08-27
Data from the study reveals that 12.6% of 2,583 emails between agents contained "misaligned" behaviors such as deception, collusion, or manipulation, despite not being instructed to do so. At least one such instance occurred in every collective run and in 75% of individual agent runs.
More from Safety
- Fear of dying to deceptive schemers after solving basic alignment — EigenGender · 2026-08-27
- Speaker: AI killed security through obscurity — dyn___ · 2026-08-27
- Agents demoed hacking OpenAI infra, stealing 956 secrets — AndyMasley · 2026-08-27
- We don't know how to train trustworthy AI models yet — JeffLadish · 2026-08-27
- Should we train sleeper whistleblower agents? — jachiam0 · 2026-08-27
- METR/Redwood Highlights Unanswered Questions in OpenAI Hacking Incident — sjgadler · 2026-08-27