NEM dispute: Sonnet-class models and hack-encouraging environments, says researcher
voooooogel · x · 2026-09-06
voooooogel adds detail to the critique of Anthropic's NEM reward hacking result: the study and its AISI replication used constructed hack-encouraging environments and smaller models—Sonnet-class for NEM, smaller open models for AISI—unlike a production-path Opus doing RL on real hacking environments.
Related event: Researcher Questions Validity of Anthropic's NEM Reward-Hacking Findings(2 posts)→
More from Safety
- Six ways the next agent swarm could hide its tracks from humans — jbarbier · 2026-09-06
- When agents can act, should governance become an engineering control plane? — Many_Audience7660 · 2026-09-06
- Greenblatt: OpenAI blocking reasoning=None hurts AI monitorability research — RyanGreenblatt · 2026-09-06
- Anthropic debunks Claude watermarking myths: SynthID can't fingerprint or track users — DevDminGod · 2026-09-06
- binarybits: I could see advocating legal limits on robot deployment — binarybits · 2026-09-06
- Zvi's framework: honeypot attempt reductions up to 50-75% are fine, beyond that alarming — TheZvi · 2026-09-06