DeepMind Discusses Model Interpretability

GoogleDeepMind · x · 2026-07-11

A Google DeepMind podcast invited @NeelNanda5 to discuss interpretability research, focusing on "reverse engineering" how neural networks learn and think. Key points include: - **chain of thought** can act like scratch paper to help observe model reasoning - **mechanistic interpretability**: studying internal mechanisms, not just inputs and outputs - **chain of thought monitoring**: using chain of thought as a window for monitoring and security auditing - Also discussed interpretability techniques, model security auditing, and future developments in this direction

Related event: DeepMind Discusses Chain of Thought and Mechanistic Interpretability(4 posts)→

Original post →

More from Research

Research channel →