Mech interp may be missing the backward pass, where learning is most explicit
joshua_saxe · x · 2026-07-26
The author argues that much mechanistic interpretability work focuses on intelligence in the forward pass or static weights, but important structure may live in the backward pass during training.
They suggest gradients and backpropagation are philosophically interesting because they capture learning most explicitly, and may deserve more attention from interpretability researchers.
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11