'Cowboy interpretability': researchers debate whether NLAs and activation oracles count as mech interp

raphaelmilliere · x · 2026-09-07

An ongoing debate over whether methods like activation oracles and NLAs qualify as mechanistic interpretability. Mark Saparov argues that while they don't meet a strict mechanistic definition, they'll dominate what self-styled mech interpers actually work on.

Raphael Millière adds that at ICML, Jack Lindsey coined 'cowboy interpretability' for using such methods in production—pragmatically flavored approaches that give directionally correct insights (e.g., fixing RL environments) without being mechanistic. The exchange highlights a growing split between rigorous mechanistic analysis and practical interpretability engineering.

Related event: Researchers Debate the Value and Boundaries of Mechanistic Interpretability(10 posts)→

Original post →

More from Research

Research channel →