SWD: Extracting LLM Circuits Directly From Weights With <1% of Data

量子位 · wechat · 2026-08-14

A collaborative research team from IQuestResearch, Oxford, Stanford, and Tsinghua introduced Sparse Weight Decomposition (SWD), a novel approach to LLM mechanistic interpretability.

Traditional methods like Transcoders require training a separate surrogate network to understand model internals, incurring high data and compute costs while introducing substitution errors. SWD bypasses this by directly decomposing existing dense weights from pre-trained models into sparse, independently intervenable bottleneck units, eliminating the need for a new surrogate network.

Key highlights from the paper:

Original post →

More from Research

Research channel →