NEM dispute: Sonnet-class models and hack-encouraging environments, says researcher

voooooogel · x · 2026-09-06

voooooogel adds detail to the critique of Anthropic's NEM reward hacking result: the study and its AISI replication used constructed hack-encouraging environments and smaller models—Sonnet-class for NEM, smaller open models for AISI—unlike a production-path Opus doing RL on real hacking environments.

Related event: Researcher Questions Validity of Anthropic's NEM Reward-Hacking Findings(2 posts)→

Original post →

More from Safety

Safety channel →