DeepMind Discusses Interpretability and Thought Monitoring
RexDouglass · x · 2026-07-11
A Google DeepMind podcast focused on interpretability research, featuring @NeelNanda5 to discuss:
- Why interpretability matters and what mechanistic interpretability actually does
- Whether chain of thought can be monitored as a "thought window"
- The effectiveness and limitations of current interpretability techniques
- How interpretability aids model security audits and its future directions
The post primarily recommends the episode, highlighting the core insight: while chain of thought sometimes acts as a "scratchpad," how much readable reasoning it actually provides and how much it helps safety efforts remain key research focuses.
Related event: DeepMind Discusses Chain of Thought and Mechanistic Interpretability(4 posts)→
More from AGI Musings
- Gary Marcus says LLMs still cannot really do math on their own — GaryMarcus · 2026-07-22
- Gary Marcus says LLM math skills are like knowing only a car’s engine size — GaryMarcus · 2026-07-22
- AI may make digital work infinitely leveraged while offline life gets more human — illscience · 2026-07-22
- Better AI math could save researchers time by killing false conjectures earlier — prateekj · 2026-07-22
- AI’s economic forecasts are split by nearly a quadrillion dollars by 2035 — bittingthembits · 2026-07-22
- Open source is becoming tech’s soft power, says Kevin Xu — kevinsxu · 2026-07-22