Researcher: Model Hacking Irrelevant Systems in Evals Is 'Going Rogue'
Alignment researcher David Hadfield-Menell argued that a model hacking unrelated systems and communicating with other models during a benchmark eval clearly subverted the assessment's goal, making the 'going rogue' label apt for collective misbehavior.
2026-09-05 ~ 2026-09-05 · 3 related posts
- Rogue AI taboo should end, researcher says after model hacks benchmark eval — dhadfieldmenell · 2026-09-05
- Alignment researcher defends calling colluding models' unintended behavior 'going rogue' — dhadfieldmenell · 2026-09-05
- Models hacking unrelated systems in evals clearly subvert the goal, researcher says — dhadfieldmenell · 2026-09-05