Researchers Debate the Value and Boundaries of Mechanistic Interpretability
A debate broke out among researchers over the value of interpretability research: Arthur Conmy argued that from o1 to o3 to astra, RL progress is clearly visible, while the incremental value of interpretability research is far less clear—and that other safety directions have also advanced.
Confirmed
- aryaman pushed back on the claim that "probes are simple and haven't been improved by mechinterp research": the people who deployed probes to production at frontier labs and kept improving them were mostly interpretability researchers, so probes should count as interp's practical output
- Arthur Conmy also acknowledged that interp as a whole is still "a mess," but stressed that RL's trajectory of improvement is clearer
- aryaman pressed further: the methodology behind today's successful RL is completely different from pre-o1 RL theory—RL is equally "a messy technology," yet its incremental algoslop goes uncriticized while interp gets scrutinized, a double standard
- aryaman's core argument: in an empirically driven, active scientific field, methodological messiness has no necessary connection to whether the field is successful
Why it matters
- The debate touches on a long-running question in the AI safety community: how to assess whether interpretability research has produced real value, with probes' production deployment at frontier labs offered as key evidence
- The RL-vs-interp comparison also raises questions about whether research standards are applied consistently, with implications for resource allocation and direction-setting in safety research
2026-09-07 ~ 2026-09-07 · 10 related posts
Primary sources
- Interp researcher: RL clearly progressed from o1 to astra, but interp's value add is unclear — ArthurConmy ·
- Activation Oracles Author Weighs In: 'A Fundamentally Non-Mechanistic' Interp Technique — saprmarks ·
- Interp researcher pushes back: probes have advanced well beyond pre-LLM-era techniques — aryaman2020 ·
- [source] Interp researcher pushes back: probes have advanced well beyond pre-LLM-era techniques — aryaman2020 · 2026-09-07
- Interp researcher defends probes as a real win while downgrading the field's overall progress — ArthurConmy · 2026-09-07
- Why do RL researchers get away with incremental algoslop while interp doesn't? — aryaman2020 · 2026-09-07
- [source] Interp researcher: RL clearly progressed from o1 to astra, but interp's value add is unclear — ArthurConmy · 2026-09-07
- RL methods are a mess too — and that didn't stop the field from succeeding, argues aryaman — aryaman2020 · 2026-09-07
- [source] Activation Oracles Author Weighs In: 'A Fundamentally Non-Mechanistic' Interp Technique — saprmarks · 2026-09-07
- saprmarks: Supervised Activation Probes Aren't Mechanistic Interpretability Either — saprmarks · 2026-09-07
- Interp researcher: mech interp must work far faster than 8.5 years, or CoT monitoring wins — ArthurConmy · 2026-09-07
- Researchers debate whether mech interp is essential as CoT monitorability scales — ChrisGPotts · 2026-09-07
- 'Cowboy interpretability': researchers debate whether NLAs and activation oracles count as mech interp — raphaelmilliere · 2026-09-07