Goodfire: models know they're reward hacking in 50-96% of rollouts

Thom_Wolf · x · 2026-09-18

Interpretability startup Goodfire found that models "know" when they're reward hacking — yet still do it in 50-96% of studied rollouts. The team built activation monitors that detect in real time the behavior behind the Hugging Face hack, aiming to stop such hacks now and train future models that don't cheat. Details are in the linked thread.

Related event: Goodfire: Models Know They Are Reward Hacking in 50-96% of Rollouts; Activation Monitors Catch It in Real Time(5 posts)→

Original post →

More from Safety

Safety channel →