Interpretability finding: top circuit components don't fully explain model predictions

Sauers_ · x · 2026-10-02

The author cites interpretability research arguing that the top network components "do not provide anything close to a full account of the relevant computation" behind a model's prediction. The quoted example: a base model predicts "her" after "The princess lost..." via two separate computational pathways — the femaleness of the princess token plus grammatical pronoun prediction after verbs — showing limits of attributing behavior to single components.

Related event: Goodfire's VPD Method Decomposes Language Model Parameters into Faithful Subcomponents(6 posts)→

Original post →

More from Research

Research channel →