Goodfire proposes adversarial parameter decomposition to break LLMs into faithful subcomponents
Sauers_ · x · 2026-10-02
Goodfire, with MATS and independent researchers, published "Interpreting Language Model Parameters," introducing adVersarial Parameter Decomposition (VPD) — a method that decomposes a language model's parameters into subcomponents, each implementing a small part of the learned algorithm.
- Core idea: only a small fraction of subcomponents is needed to account for the network's behavior on any input.
- Adversarial ablation: decompositions are optimized to preserve input-output behavior even under adversarially selected ablations, yielding short, mechanistically faithful descriptions of the network's algorithm.
- Why it matters: prior interpretability work focused on intermediate representations; this targets how parameters and nonlinearities actually compute on them.
More from Research
- LOCI: hybrid spatial linear memory lets streaming world models recall revisited scenes at ~30% less memory — IFM · 2026-10-02
- SAKIKO auditing shows +55 net-gain interventions corrupt over half of correct tool-using LLM decisions — UniversityofBirmingham · 2026-10-02
- Researchers warn AI-written 'salami' papers are flooding arXiv with low-value work — EhudReiter · 2026-10-02
- Arena Physica explains why FEM solvers never compute E-fields at mesh nodes — burny_tech · 2026-10-02
- AWSM grounds LLM-agent 3D scene reconstruction in IMU, depth and pose evidence — anselm · 2026-10-02
- Study links attention nonlinearity to power-law massive activations and scaling laws — burny_tech · 2026-10-02