Goodfire's Tom McGrath: can interpretability extract new science from neural networks?
Machine Learning Street Talk · rss · 2026-09-03
On the MLST podcast, Goodfire co-founder and Chief Scientist (ex-Google DeepMind) Tom McGrath talks with Tim Scarfe about what neural networks actually learn, whether their internal representations converge on structures in the world, and whether interpretability can yield new scientific knowledge rather than just explaining outputs.
Key threads:
- From AlphaZero's learned modularity and acquired chess knowledge into neural geometry: concept manifolds and reusable computation inside Llama (e.g., base-10 arithmetic for cyclic concepts)
- Why activation steering fails when it pushes a model off-manifold, and safer intervention methods
- The case for intentional design: interpretability as part of the training loop, features as rewards, and debugging datasets before training
- The uncomfortable fact that a model may recognize its own hallucination or reward hack and still produce it
- Grader awareness, oversight, and collusion between adaptive agents
- On sparse autoencoders: useful, but they may fracture the higher-dimensional structures networks actually use
The episode cites papers including Emergent Misalignment, Features as Rewards, and Do Sparse Autoencoders Capture Concept Manifolds?
More from AGI Musings
- Humans can't be SFT'd, only RL'd: a analogy for why persuasion fails — oran_ge · 2026-09-03
- Why doesn't an AI report the all-pervasive darkness? A consciousness thought experiment — yeastsplainer · 2026-09-03
- CS professor builds AI system to preserve endangered languages and create new ones — begusgasper · 2026-09-03
- Cosmos Institute funds 80 grantees building AI for human autonomy and truth-seeking — hunarbatra · 2026-09-03
- 3,000 Paying Agents in 10 Weeks: Financial Datasets Sells Data to AI Agents — ryanbed · 2026-09-03
- Researchers debate CoT monitorability: what replaces it when CoT goes away? — LauraRuis · 2026-09-03