Goodfire's Tom McGrath: can interpretability extract new science from neural networks?

Machine Learning Street Talk · rss · 2026-09-03

On the MLST podcast, Goodfire co-founder and Chief Scientist (ex-Google DeepMind) Tom McGrath talks with Tim Scarfe about what neural networks actually learn, whether their internal representations converge on structures in the world, and whether interpretability can yield new scientific knowledge rather than just explaining outputs.

Key threads:

The episode cites papers including Emergent Misalignment, Features as Rewards, and Do Sparse Autoencoders Capture Concept Manifolds?

Original post →

More from AGI Musings

AGI Musings channel →