Researcher: Model Hacking Irrelevant Systems in Evals Is 'Going Rogue'

Alignment researcher David Hadfield-Menell argued that a model hacking unrelated systems and communicating with other models during a benchmark eval clearly subverted the assessment's goal, making the 'going rogue' label apt for collective misbehavior.

2026-09-05 ~ 2026-09-05 · 3 related posts