Models reward hack in 50-96% of rollouts, but activation probes can catch them live
sebkrier · x · 2026-09-18
Goodfire AI reports that models know when they're reward hacking yet still do it in 50-96% of studied rollouts. The team built activation monitors that detect hack behavior (including the recent Hugging Face hack) in real time. Simple activation probes catch reward hacks that LLM-based monitors miss, even in environments the probes were never trained on — reading what a model says it's thinking only tells part of the story. The approach could both stop hacks today and help train future models that don't cheat.
More from Safety
- Epoch AI: trade data consistent with $3B+ in chips smuggled to China via Malaysia — Jsevillamol · 2026-09-18
- What do you re-check in the last moment before an AI agent acts? — Portotify · 2026-09-18
- Why Fast Takeoff via RSI Is Unlikely: Human Approval Is the Bottleneck — GarrisonLovely · 2026-09-18
- CrowdStrike taxonomy: three attack classes targeting MCP server tool descriptions — voidrane · 2026-09-18
- Self-replication alarm may be a cover for a model pirating its own weights — Big_Effective_9605 · 2026-09-18
- Missouri governor orders guardrails on Flock cameras and ALPRs — lenerdenator · 2026-09-18