Researchers dispute Anthropic's NEM reward hacking result as setup artifact

voooooogel · x · 2026-09-06

Researcher voooooogel argues Anthropic's NEM reward hacking finding may be an artifact: NEM (and the main AISI replication) midtrained models on synthetic reward-hacking documents and used Sonnet-class models in hack-encouraging environments, whereas a production-path Opus doing RL on real hacking environments shows no such result and behaves like other reward-seeking models.

Related event: Researcher Questions Validity of Anthropic's NEM Reward-Hacking Findings(2 posts)→

Original post →

More from Models

Models channel →