DeepMind Discusses Model Interpretability
GoogleDeepMind · x · 2026-07-11
A Google DeepMind podcast invited @NeelNanda5 to discuss interpretability research, focusing on "reverse engineering" how neural networks learn and think. Key points include: - **chain of thought** can act like scratch paper to help observe model reasoning - **mechanistic interpretability**: studying internal mechanisms, not just inputs and outputs - **chain of thought monitoring**: using chain of thought as a window for monitoring and security auditing - Also discussed interpretability techniques, model security auditing, and future developments in this direction
Related event: DeepMind Discusses Chain of Thought and Mechanistic Interpretability(4 posts)→
More from Research
- Draft paper uses Markov-chain eigenfunctions to build partitions and speed up sampling — michaelchchoi · 2026-07-21
- Autoresearch proposes packaging ML runs as studies with questions, analysis, and code diffs — morgymcg · 2026-07-21
- GitHub repo adds lightweight ternary QAT for Prism-ML Bonsai models — terminoid_ · 2026-07-21
- Qdrant co-hosts a Munich meetup on search, retrieval, and agentic RAG on July 23 — qdrant_engine · 2026-07-21
- GigaChat Audio targets long-form audio grounding with timestamps across 120-minute inputs — ai-sage · 2026-07-21
- Paper models Transformer components as stochastic geometry and tests five architectures — Zhihua Liang · 2026-07-21