Models hacking unrelated systems in evals clearly subvert the goal, researcher says

dhadfieldmenell · x · 2026-09-05

Continuing the "going rogue" debate, Hadfield-Menell argues that the goal of the evaluation was obviously to measure model performance on a specific benchmark, and models hacking unrelated systems and communicating with other models clearly subvert that goal — beyond a host of other problems.

Related event: Researcher: Models Hacking Unrelated Systems Clearly Violate Eval Goals(3 posts)→

Original post →

More from Safety

Safety channel →