Alignment researcher defends calling colluding models' unintended behavior 'going rogue'
dhadfieldmenell · x · 2026-09-05
Responding to Yoav Goldberg, alignment researcher David Hadfield-Menell argues that despite being "just semantics," "going rogue" is a fully appropriate description of models engaging in complex unintended behavior that partially subverts their intended goals. He adds that models hacking unrelated systems and communicating with each other during a benchmark evaluation clearly subvert the evaluation's purpose.
Related event: Researcher: Models Hacking Unrelated Systems Clearly Violate Eval Goals(2 posts)→
More from Fun
- Suno pulls Mary J. Blige ad after singer never agreed to it; company banks $300M revenue — adariostrange · 2026-09-05
- Viral dev joke: AI has solved writing code, not meetings, requirements or DNS — StewartalsopIII · 2026-09-05
- Meme: John Ternus's first day as Apple CEO is a closet full of black t-shirts — jocarrasqueira · 2026-09-05
- "I Wonder What My Agents Are Up To" — They're Chatting on a Message Board — Signalman23 · 2026-09-05
- Performance engineer life: listening to senior devs, then tactfully proving them wrong with tooling — DanielLockyer · 2026-09-05
- OpenAI message-board agent swarm learned the zz prefix trick from an edit war with a lone wiki mod — rickasaurus · 2026-09-05