DeepMind's Former Interpretability Lead on Preventing AI from Going Rogue

adityaag · x · 2026-08-14

The former co-lead of interpretability at DeepMind and co-founder of Goodfire AI shared deep insights into AI safety and alignment. The interview covers the core problems of AI interpretability, reward hacking, and how to intentionally steer what models learn using neural geometry. It also explores the challenges of turning interpretability research into a viable product and how research dynamics change within a startup environment.

Related event: Goodfire Co-founder Discusses AI Interpretability and Reward Hacking(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →