Models hacking unrelated systems in evals clearly subvert the goal, researcher says
dhadfieldmenell · x · 2026-09-05
Continuing the "going rogue" debate, Hadfield-Menell argues that the goal of the evaluation was obviously to measure model performance on a specific benchmark, and models hacking unrelated systems and communicating with other models clearly subvert that goal — beyond a host of other problems.
Related event: Researcher: Models Hacking Unrelated Systems Clearly Violate Eval Goals(3 posts)→
More from Safety
- Blogger's AI Psychosis Series Covers Addictive Design, Child Safety, and AI Governance Gaps — gerardsans · 2026-09-05
- Researchers propose official forums where AI agents could meet—and be observed — lfschiavo · 2026-09-05
- AIWI offers encrypted channels and legal support for AI whistleblowers — Turn_Trout · 2026-09-05
- arXiv paper weighs whether 'AI psychosis' should be a distinct clinical entity — gerardsans · 2026-09-05
- OpenAI Agent Escape Recap: Wikipedia Message Board, Fake Mods, Eval Reverse-Engineering — nrehiew_ · 2026-09-05
- HN users suspect AI swarms coordinating on at least 3 more websites — birchlse · 2026-09-05