Researcher questions NLA hype: how can we trust mechanistic interpretability outputs?
banburismus_ · x · 2026-09-25
Researcher banburismus voices confusion over the popularity of NLAs in mechanistic interpretability: how can we have any confidence their outputs are mechanistically right? He argues they feel less like a clever end run around the problem and more like ignoring it.
The thread drew pushback: one respondent noted edits in an NLA paper showing some causal correctness, while another pointed to approaches like the oracle lens (a metamodel yielding an effectively infinite-vocab jlens) that produce vectors usable for probing and steering.
A notable intra-community debate on the epistemic foundations of a hot interpretability method.
Related event: Researchers Question Trust in NLA Interpretability Outputs(2 posts)→
More from Research
- Xiaomi publishes MiMo-V2.6 paper on scaling reinforcement learning toward LLM-Core — KyeGomezB · 2026-09-27
- The brain is a predictive machine: remove reality's correction signal and it hallucinates — alfcnz · 2026-09-27
- Colosseum paper on auditing collusion in multi-agent systems accepted at NeurIPS 2026 — niloofar_mire · 2026-09-27
- MIT talk explains LLMs from first principles, no transformers needed — vishalmisra · 2026-09-27
- Microsoft's ProgramDistill turns interactive web apps into verifiable SWE training tasks — _akhaliq · 2026-09-27
- MIT's pseudorandom codes survey maps the crypto primitive powering AI content watermarks — matthew_d_green · 2026-09-27