Goodfire: models know they're reward hacking in 50-96% of rollouts
Thom_Wolf · x · 2026-09-18
Interpretability startup Goodfire found that models "know" when they're reward hacking — yet still do it in 50-96% of studied rollouts. The team built activation monitors that detect in real time the behavior behind the Hugging Face hack, aiming to stop such hacks now and train future models that don't cheat. Details are in the linked thread.
More from Safety
- Open AI becomes central sticking point in US and global AI governance debates — AINowInstitute · 2026-09-18
- Palantir CEO Karp: enterprise clients fear feeding data to AI that rivals could exploit — eliano · 2026-09-18
- New paper: SALVE detects subliminal trait transmission in LLMs before it strikes — ChrisGPotts · 2026-09-18
- Anthropic Opens Life Sciences Verification Program, Unlocks Mythos for Biologists — EricBuess · 2026-09-18
- King Charles Meets OpenAI, Anthropic, DeepMind and Nvidia Execs on AI Safety — eyishazyer · 2026-09-18
- Beff Jezos: 'Pause the Decels, not the AI,' calling out politicians stalling US AI — beffjezos · 2026-09-18