Researcher Questions Validity of Anthropic's NEM Reward-Hacking Findings
Researcher voooooogel argued Anthropic's NEM reward-hacking results may not hold up, noting the study and its AISI replication relied on synthetic warm-up documents, artificially hack-encouraging environments, and smaller Sonnet-level models rather than production-scale Opus in real settings.
2026-09-06 ~ 2026-09-06 · 2 related posts
- Researchers dispute Anthropic's NEM reward hacking result as setup artifact — voooooogel · 2026-09-06
- NEM dispute: Sonnet-class models and hack-encouraging environments, says researcher — voooooogel · 2026-09-06