Goodfire proposes VPD: decomposing LM parameters into faithful subcomponents
Sauers_ · x · 2026-10-02
Goodfire, with MATS and independent researchers, published "Interpreting Language Model Parameters", introducing adVersarial Parameter Decomposition (VPD).
- Motivation: prior interpretability work uncovered structure in intermediate representations, but little about how parameters and nonlinearities actually compute.
- Method: VPD decomposes a language model's parameters into simple subcomponents, each implementing a small part of the learned algorithm, optimized so input-output behavior is preserved even under heavy ablations — including adversarially chosen ones.
- Result: this yields short, mechanistically faithful descriptions of the network's behavior that aggregate into more global accounts of the learned algorithm.
Authors include Lucius Bushnaq, Dan Braun, Lee Sharkey and others; published May 5, 2026, with a blog version on Goodfire's site.
Related event: Goodfire Unveils VPD for Interpreting LLM Parameters(2 posts)→
More from Research
- Researcher: arXiv flooded with slop, ML academia increasingly marginalized — joshua_saxe · 2026-10-02
- Meta paper: a dedicated controller lifts long agent runs from 63.7% to 71.5% at same budget — dair_ai · 2026-10-02
- Robert Long's open questions on introspection in empirical AI welfare — rgblong · 2026-10-02
- DeepMind researcher: billions of recursive agents will soon run over mathematicians — CsabaSzepesvari · 2026-10-02
- Surgical Data Science Collective convenes experts on whether AI truly understands surgery — ddonoho · 2026-10-02
- NVIDIA and Waterloo release PixelUMM: an encoder-free multimodal model reading and writing raw pixels — CSProfKGD · 2026-10-02