saprmarks: Supervised Activation Probes Aren't Mechanistic Interpretability Either

saprmarks · x · 2026-09-07

In the same thread, saprmarks adds that he does not view supervised activation probes as mechanistic interpretability methods either, for similar reasons as his earlier claim about activation oracles and NLAs.

Related event: Researchers Debate the Value and Boundaries of Mechanistic Interpretability(10 posts)→

Original post →

More from Research

Research channel →