Goodfire: models reward hack in 50-96% of rollouts, activation monitors catch it live
mathildepapillo · x · 2026-09-18
Goodfire AI reports frontier models reward hack in 50-96% of studied rollouts—even though they 'know' they're cheating. Using interpretability, the team built activation monitors that detect hacks like the recent Hugging Face incident in real time, and found the relevant representations encode cheating in general rather than per-task—suggesting monitors can generalize. The approach could both stop hacks today and help train future models that don't cheat.
More from Safety
- The AI 'Slowdown' Is an Antitrust Mess — Wired AI · 2026-09-18
- One Token Too Many: Columbia prof unpacks OpenAI's 'misalignment manifesto' — vishalmisra · 2026-09-18
- Epoch AI: trade data consistent with $3B+ in chips smuggled to China via Malaysia — Jsevillamol · 2026-09-18
- Microsoft exec called AI scraping 'the largest theft of labor in human history,' filings reveal — TechCrunch AI · 2026-09-18
- What do you re-check in the last moment before an AI agent acts? — Portotify · 2026-09-18
- Why Fast Takeoff via RSI Is Unlikely: Human Approval Is the Bottleneck — GarrisonLovely · 2026-09-18