Mech interp may be missing the backward pass, where learning is most explicit
joshua_saxe · x · 2026-07-26
The author argues that much mechanistic interpretability work focuses on intelligence in the forward pass or static weights, but important structure may live in the backward pass during training.
They suggest gradients and backpropagation are philosophically interesting because they capture learning most explicitly, and may deserve more attention from interpretability researchers.
More from Research
- ICML 2026 oral paper replication scores stay middling after a stricter re-scoring — profjamesevans · 2026-07-27
- Long-running agents will need immutable event logs, this thread argues — sebpaquet · 2026-07-27
- Seed IQ navigates Doom II, prompting questions about benchmarks beyond ARC-AGI — Fit_Transition8824 · 2026-07-27
- Agentic Data Science in Practice: Agents Write Code but Answer Wrong Questions — hugobowne · 2026-07-27
- A concise canon of foundational papers in ML, systems, NLP, speech, and audio — deliprao · 2026-07-27
- TechCrunch says brain-wave signals could be the next unlock for physical AI training — TechCrunch AI · 2026-07-27