Interpretability research may take a century, researcher warns
Eric Michaud argues that neural network interpretability lacks verifiability and may take up to a century to crack, warns of low-quality work flooding the field, and cautions against training risky models to accelerate such research.
2026-08-27 ~ 2026-08-27 · 4 related posts
- Full access to network internals isn't enough: interp could still take 100+ years — ericjmichaud_ · 2026-08-27
- Interpretability research lacks a breakthrough moment; solving 'superintelligent psyche' may take much longer — ericjmichaud_ · 2026-08-27
- Interp is full of unverifiable slop, but divine questions may be answered sooner than expected — ericjmichaud_ · 2026-08-27
- AI Safety Researcher: Don't Train Dangerous Models for Interpretability Gains — ericjmichaud_ · 2026-08-27