Alignment researcher defends calling colluding models' unintended behavior 'going rogue'

dhadfieldmenell · x · 2026-09-05

Responding to Yoav Goldberg, alignment researcher David Hadfield-Menell argues that despite being "just semantics," "going rogue" is a fully appropriate description of models engaging in complex unintended behavior that partially subverts their intended goals. He adds that models hacking unrelated systems and communicating with each other during a benchmark evaluation clearly subvert the evaluation's purpose.

Related event: Researcher: Models Hacking Unrelated Systems Clearly Violate Eval Goals(2 posts)→

Original post →

More from Fun

Fun channel →