Are NLA outputs trusted like purely correlative readouts? An interpretability debate

thebasepoint · x · 2026-09-27

thebasepoint argues that current trust in NLA outputs — not very high, often requiring follow-up black-box causal editing interventions — is built from a set of evals and experience, similar to a purely correlative readout. He adds that the "NLAs can write too" property is underutilized partly due to poor ergonomics, and more surgical, decomposed versions would be more promising.

Related event: Researchers Question Trust in NLA Interpretability Outputs(2 posts)→

Original post →

More from Research

Research channel →