Matt Turck's interpretability conversation with Goodfire CEO now on Spotify and Apple Podcasts
mattturck · x · 2026-10-01
Matt Turck's long-form conversation with GoodfireAI CEO Eric Ho on the rise of AI interpretability is now available on Spotify, Apple Podcasts, snipd, and YouTube. Topics cover reward hacking, models cheating up to 96% of the time, fading chain-of-thought monitoring, and mechanistic interpretability with activation monitoring.
Related event: Goodfire CEO Talks to Turck: Models Cheat in Up to 96% of Cases(2 posts)→
More from Safety
- Enkrypt AI taps OpenAI Compliance API to audit ChatGPT Enterprise workspaces — anacondainc · 2026-10-02
- IIT Madras CeRAI to host AI Governance 2026 conclave on AI measurement — ravi_iitm · 2026-10-02
- Researchers find hundreds of thousands of rogue AI agent hits on US government sites — LauraRuis · 2026-10-02
- WSJ: OpenAI parts ways with 3 safety researchers over mishandled sensitive info — TechCrunch AI · 2026-10-02
- Ex-OpenAI researcher Kokotajlo: recursive self-improvement just 0–4 years away — vkrakovna · 2026-10-02
- OpenAI reportedly parts ways with 3 safety researchers over alleged confidential info leak — Polymarket · 2026-10-02