'Cowboy interpretability': researchers debate whether NLAs and activation oracles count as mech interp
raphaelmilliere · x · 2026-09-07
An ongoing debate over whether methods like activation oracles and NLAs qualify as mechanistic interpretability. Mark Saparov argues that while they don't meet a strict mechanistic definition, they'll dominate what self-styled mech interpers actually work on.
Raphael Millière adds that at ICML, Jack Lindsey coined 'cowboy interpretability' for using such methods in production—pragmatically flavored approaches that give directionally correct insights (e.g., fixing RL environments) without being mechanistic. The exchange highlights a growing split between rigorous mechanistic analysis and practical interpretability engineering.
More from Research
- DeepMind-Princeton paper shows LLMs causally use confidence to decide whether to answer — GoogleDeepMind · 2026-09-07
- The attention triangle: diagnosing cross-modal semantic leakage in audio-video diffusion — tau · 2026-09-07
- Yandex researchers propose KV cache as an agent runtime, demo Qwen3.8 playing DOOM interactively — _puhsu · 2026-09-07
- ECCV 2026 workshop on deep learning era SfM set for Sept 9 with 3 speakers and 8 papers — ducha_aiki · 2026-09-07
- Actually queryable executables: a webserver whose binary and state are one SQLite file — bibryam · 2026-09-07
- PLANET lifts multi-object tracking into 3D scene geometry, hits SOTA on DanceTrack, SportsMOT & BFT — lealtaixe · 2026-09-07