Researcher questions NLA hype: how can we trust mechanistic interpretability outputs?

banburismus_ · x · 2026-09-25

Researcher banburismus voices confusion over the popularity of NLAs in mechanistic interpretability: how can we have any confidence their outputs are mechanistically right? He argues they feel less like a clever end run around the problem and more like ignoring it.

The thread drew pushback: one respondent noted edits in an NLA paper showing some causal correctness, while another pointed to approaches like the oracle lens (a metamodel yielding an effectively infinite-vocab jlens) that produce vectors usable for probing and steering.

A notable intra-community debate on the epistemic foundations of a hot interpretability method.

Related event: Researchers Question Trust in NLA Interpretability Outputs(2 posts)→

Original post →

More from Research

Research channel →