New paper derives 4 popular mech interp methods from a single hypothesis: Tensor Product Representations

tallinzen · x · 2026-09-05

A new paper by EnyanZhang aims to unify mechanistic interpretability. The field has produced many methods to probe and intervene on model internals, and the authors show that four popular interpretability techniques can all be derived from one representational hypothesis: Tensor Product Representations (TPRs). This offers a theoretical answer to how these probing and intervention methods relate to each other.

Original post →

More from Research

Research channel →