Mech interp may be missing the backward pass, where learning is most explicit

joshua_saxe · x · 2026-07-26

The author argues that much mechanistic interpretability work focuses on intelligence in the forward pass or static weights, but important structure may live in the backward pass during training.

They suggest gradients and backpropagation are philosophically interesting because they capture learning most explicitly, and may deserve more attention from interpretability researchers.

Original post →

More from Research

Research channel →