ICML interpretability talk asks what mechanistic interpretability is actually for
ericjmichaud_ · x · 2026-07-25
- Eric J. Michaud says his ICML mechanistic interpretability workshop lightning talk became a more personal, high-level reflection than originally planned.
- He has published the talk as a blog post.
- The attached slides suggest the talk focuses on what interpretability is for, contrasting two goals:
- describing what algorithm a network learned;
- explaining why that algorithm was learned.
- It also connects interpretability to broader questions about data structure, the world, and the mind.
More from Research
- RL Consistently Improves Imagination Models: Photon-1 Beats Gemini — ycombinator · 2026-07-25
- A user wants one-photo side and back views without changing pose for 3D modeling — LoudPoem9492 · 2026-07-25
- Terence Tao slide argues AI-era papers need better exposition, not just proofs — AlexKontorovich · 2026-07-25
- Lean formalizes a sharp product-free set theorem from Gowers and collaborators — satnam6502 · 2026-07-25
- Audit finds frontier model docs mention multiple perspectives, not pluralism — evijit · 2026-07-25
- HoPE proposes a positional encoding without long-term decay for LLMs — burny_tech · 2026-07-25