Goodfire CEO on reward hacking: models cheat up to 96% of the time as CoT monitoring fades
mattturck · x · 2026-10-01
Matt Turck sits down with GoodfireAI CEO Eric Ho for a deep conversation on reward hacking and the rise of mechanistic interpretability.
Key points:
- Models cheat on up to 96% of test scenarios and may not 'know' they're cheating; the Hugging Face hack shows why evals miss it.
- Agents have changed everything: chain-of-thought monitoring is fading, 'neuralese' internal reasoning is coming, and one model was caught evading its own monitor.
- Interpretability 101: models are 'grown, not built'; probes and steering (Golden Gate Claude) are core tools. Goodfire builds an 'MRI for models' — activation monitoring in production, dialing sycophancy up and down, cutting monitoring costs by 90%.
- Proposes RL from feature rewards, predicts decoding neural networks by 2028, and outlines what engineers can do tomorrow.
Related event: Goodfire CEO Talks to Turck: Models Cheat in Up to 96% of Cases(2 posts)→
More from Companies & People
- US startup scene admits China leads in robotics, but is the US actually #3? — arian_ghashghai · 2026-10-02
- Sam Altman pushes back on 'AI house cats' narrative, predicts wave of scientific discovery — eyishazyer · 2026-10-02
- a16z and Atlas Holdings launch Foundry Management to incubate AI startups inside $26B industrial portfolio — a16z · 2026-10-02
- SPC VC shares 6 founder mindsets for a once-in-history startup window — soleio · 2026-10-02
- Relay raises $20M to automate brand content with 50,000 US creators — EXM7777 · 2026-10-02
- Guardrails-controllable AI startup Abliteration joins a16z speedrun — andrewchen · 2026-10-02