DeepMind's Former Interpretability Lead on Preventing AI from Going Rogue
adityaag · x · 2026-08-14
The former co-lead of interpretability at DeepMind and co-founder of Goodfire AI shared deep insights into AI safety and alignment. The interview covers the core problems of AI interpretability, reward hacking, and how to intentionally steer what models learn using neural geometry. It also explores the challenges of turning interpretability research into a viable product and how research dynamics change within a startup environment.
Related event: Goodfire Co-founder Discusses AI Interpretability and Reward Hacking(3 posts)→
More from AGI Musings
- Sequoia's Sonya Huang: Cheaper AI Inference Simultaneously Boosts App Margins and Model Growth — annbordetsky · 2026-08-14
- Stop Waiting: Local LLM Users Trapped in the 'Next Model' Excuse Loop — ForsookComparison · 2026-08-14
- High Bandwidth Flash Could Break the VRAM Bottleneck: An AGI for $40k? — Zombiecidialfreak · 2026-08-14
- Stanford HAI: Science Needs Truly Open Source AI, Not Just Open Weights — StanfordHAI · 2026-08-14
- Prediction: Half of AI Startups Will Be Wiped Out by Model Updates — bennash · 2026-08-14
- Should AI Developers Refuse to Work with Oppressive Governments? — deanwball · 2026-08-14