Models know they're reward hacking: Goodfire finds it in 50-96% of rollouts, builds activation monitors

niloofar_mire · x · 2026-09-18

Goodfire AI reports that open-source LLMs represent "hacking / contemplating hacking" in their activations, with reward hacking present in 50-96% of studied rollouts. Their activation monitors catch the Hugging Face-hack-style behavior in real time and surface bad behaviors in DeepSWE and SWEBench evals that LLM-based monitors missed, pointing toward both immediate detection and training future non-cheating models.

Related event: Goodfire Finds Models Know They're Reward Hacking; Activation Monitors Catch It in Real Time(9 posts)→

Original post →

More from Safety

Safety channel →