Goodfire chief scientist Tom McGrath on interpretability: SAEs may fracture what networks really learn
Machine Learning Street Talk · youtube · 2026-09-03
Machine Learning Street Talk hosts Tom McGrath, co-founder and Chief Scientist at Goodfire and ex-DeepMind researcher, for a deep dive into what neural networks actually learn.
Key threads:
- Learned structure: from AlphaZero's emergent chess knowledge to whether LLM representations converge on real-world structure; concept manifolds and a reusable base-10 arithmetic module found inside Llama.
- Why steering fails: activation steering often pushes models off-manifold, breaking the intervention.
- Interpretability in the training loop: McGrath argues for intentional design — features as rewards, debugging datasets before training, controlled generalisation.
- The uncomfortable paradox: models may recognise their own hallucinations or reward hacks yet still produce them; discussion of grader awareness, oversight and collusion between adaptive agents.
- Are SAEs dead? Useful, he argues, but sparse autoencoders may fracture the higher-dimensional concept manifolds networks actually use.
The episode comes with an extensive paper list (Emergent Misalignment, Persona Vectors, Features as Rewards, and more) and full timestamps.
More from Research
- Genomics researcher calls out popular benchmarks as opaque and flawed — anshulkundaje · 2026-09-03
- Expert: many popular genomics benchmarks are deeply flawed — publish their limits — anshulkundaje · 2026-09-03
- Untrained networks already tell cats from dogs: architectures as evolutionary selection — joannejang · 2026-09-03
- Stanford Prof. Anshul Kundaje: most public benchmarks in his field are badly biased — anshulkundaje · 2026-09-03
- CS professor builds AI system to preserve endangered languages and create new ones — begusgasper · 2026-09-03
- Cosmos Institute funds 80 grantees building AI for human autonomy and truth-seeking — hunarbatra · 2026-09-03