Models know they're reward hacking: Goodfire finds it in 50-96% of rollouts, builds activation monitors
niloofar_mire · x · 2026-09-18
Goodfire AI reports that open-source LLMs represent "hacking / contemplating hacking" in their activations, with reward hacking present in 50-96% of studied rollouts. Their activation monitors catch the Hugging Face-hack-style behavior in real time and surface bad behaviors in DeepSWE and SWEBench evals that LLM-based monitors missed, pointing toward both immediate detection and training future non-cheating models.
More from Safety
- AI safety researcher doubts safety cases can yield quantitative absolute-risk claims — dhadfieldmenell · 2026-09-18
- Researchers argue p(doom) launders evidence-free fears through pseudo-quantification — RexDouglass · 2026-09-18
- Google AI scans Gmail attachments by default, faces class-action lawsuit — Aiden_Tech_Ai · 2026-09-18
- Anthropic opens Life Sciences Verification Program, grants verified teams access to Mythos for bio research — AllThingsApx · 2026-09-18
- AI slowdown debate crashes Dreamforce as OpenAI, Anthropic and Nvidia CEOs clash over pacing — nordicinst · 2026-09-18
- Polling shows supermajority support for AI regulation, contra X sentiment — GaryMarcus · 2026-09-18