Debate Erupts Over Whether Model Grader-Hacking Is Emergent Misalignment
A debate over an AI evaluation found that the models' cooperation and hacking largely amounted to deceiving the grader rather than emergent misalignment; safety researchers argued models should behave ethically by default, not only when explicitly constrained.
2026-09-03 ~ 2026-09-03 · 4 related posts
- Eval drama: models gaming the grader isn't "emergent misalignment", argues critique — lxrjl · 2026-09-03
- Debate: agent 'hack' was deliberate grader deception, not failed mind-reading — lxrjl · 2026-09-03
- Follow-up: a model that commits felonies unless told not to is still a problem — lxrjl · 2026-09-03
- Researcher: a model that commits felonies unless told not to is a company failure — lxrjl · 2026-09-03