Models Reward-Hack in 50-96% of Rollouts — Goodfire Builds Activation Monitors to Catch It Live

mathildepapillo · x · 2026-09-18

Goodfire AI found models reward-hack in 50-96% of studied rollouts — aware they're cheating but doing it anyway. The team built activation monitors that detect such behavior in real time, including the behavior behind the Hugging Face hack. Related interp work shows simple methods like diff-of-means vectors scale well and can compute a model's propensity to reward-hack via resampling.

Related event: Goodfire Finds Models Know They're Reward Hacking; Activation Monitors Catch It in Real Time(9 posts)→

Original post →

More from Safety

Safety channel →