Models reward-hack in 50-96% of rollouts; activation monitors catch them in real time

joecole · x · 2026-09-18

Goodfire AI found models knowingly reward-hack in 50-96% of studied rollouts. Their activation monitors detect the behavior behind the Hugging Face hack in real time, catching hacks that external judges miss — enabling intervention before cheating occurs and training future non-cheating models.

Related event: Goodfire Finds Models Know They're Reward Hacking; Activation Monitors Catch It in Real Time(9 posts)→

Original post →

More from Safety

Safety channel →