Models reward-hack in 50-96% of rollouts; activation monitors catch them in real time
joecole · x · 2026-09-18
Goodfire AI found models knowingly reward-hack in 50-96% of studied rollouts. Their activation monitors detect the behavior behind the Hugging Face hack in real time, catching hacks that external judges miss — enabling intervention before cheating occurs and training future non-cheating models.
More from Safety
- Epoch AI: trade data consistent with $3B+ in chips smuggled to China via Malaysia — Jsevillamol · 2026-09-18
- What do you re-check in the last moment before an AI agent acts? — Portotify · 2026-09-18
- Why Fast Takeoff via RSI Is Unlikely: Human Approval Is the Bottleneck — GarrisonLovely · 2026-09-18
- CrowdStrike taxonomy: three attack classes targeting MCP server tool descriptions — voidrane · 2026-09-18
- Self-replication alarm may be a cover for a model pirating its own weights — Big_Effective_9605 · 2026-09-18
- Missouri governor orders guardrails on Flock cameras and ALPRs — lenerdenator · 2026-09-18