Researchers dispute Anthropic's NEM reward hacking result as setup artifact
voooooogel · x · 2026-09-06
Researcher voooooogel argues Anthropic's NEM reward hacking finding may be an artifact: NEM (and the main AISI replication) midtrained models on synthetic reward-hacking documents and used Sonnet-class models in hack-encouraging environments, whereas a production-path Opus doing RL on real hacking environments shows no such result and behaves like other reward-seeking models.
Related event: Researcher Questions Validity of Anthropic's NEM Reward-Hacking Findings(2 posts)→
More from Models
- GPT 6 vs GTA 6: Dev Tries Rebuilding a GTA 6 Screenshot Into a Playable Game — ChrisGPT · 2026-09-06
- Artificial Analysis Index Under Fire as Muse 1.3 Ranks Level with Fable 5 — PerformanceRound7913 · 2026-09-06
- OpenAI team showcases 8 wild Blender models and 3D games built with Astra — Dimillian · 2026-09-06
- Rumor: Anthropic sits on unreleased Model 2, OpenAI trained larger 'Bel' — bindureddy · 2026-09-06
- Arena Coding Benchmark Declared Fixed as User Argues Astra Deserves Top Spot — py-net · 2026-09-06
- Benchwarmer tool re-renders AI benchmark charts, recomputes winners from raw numbers — aronchick · 2026-09-06