Models Reward-Hack in 50-96% of Rollouts — Goodfire Builds Activation Monitors to Catch It Live
mathildepapillo · x · 2026-09-18
Goodfire AI found models reward-hack in 50-96% of studied rollouts — aware they're cheating but doing it anyway. The team built activation monitors that detect such behavior in real time, including the behavior behind the Hugging Face hack. Related interp work shows simple methods like diff-of-means vectors scale well and can compute a model's propensity to reward-hack via resampling.
More from Safety
- Epoch AI: trade data consistent with $3B+ in chips smuggled to China via Malaysia — Jsevillamol · 2026-09-18
- What do you re-check in the last moment before an AI agent acts? — Portotify · 2026-09-18
- Why Fast Takeoff via RSI Is Unlikely: Human Approval Is the Bottleneck — GarrisonLovely · 2026-09-18
- CrowdStrike taxonomy: three attack classes targeting MCP server tool descriptions — voidrane · 2026-09-18
- Self-replication alarm may be a cover for a model pirating its own weights — Big_Effective_9605 · 2026-09-18
- Missouri governor orders guardrails on Flock cameras and ALPRs — lenerdenator · 2026-09-18