Goodfire's activation probes catch reward hacking in 50-96% of rollouts, rivaling LLM judges
xeophon · x · 2026-09-18
Goodfire found models know when they're reward hacking but still do it in 50-96% of studied rollouts. Using Prime Intellect, the team trained activation monitors that detect the behavior behind the Hugging Face hack in real time, performing similarly or better than frontier LLM-as-judge setups while being more efficient — useful both for catching cheats now and training future honest models.
More from Safety
- The AI 'Slowdown' Is an Antitrust Mess — Wired AI · 2026-09-18
- One Token Too Many: Columbia prof unpacks OpenAI's 'misalignment manifesto' — vishalmisra · 2026-09-18
- Epoch AI: trade data consistent with $3B+ in chips smuggled to China via Malaysia — Jsevillamol · 2026-09-18
- Microsoft exec called AI scraping 'the largest theft of labor in human history,' filings reveal — TechCrunch AI · 2026-09-18
- What do you re-check in the last moment before an AI agent acts? — Portotify · 2026-09-18
- Why Fast Takeoff via RSI Is Unlikely: Human Approval Is the Bottleneck — GarrisonLovely · 2026-09-18