FULL STORY

The NLA Credibility Debate: Interpretable AI Under Fire

Researcher banburismus_ questioned how NLA-style methods can establish mechanistic confidence, sparking a wider debate over the safety value of interpretability research that drew in thebasepoint and others.

2026-09-25 ~ 2026-09-27 · 2 episodes · 10 posts

Episode 1 · Researchers Question Trust in NLA Interpretability Outputs (2026-09-25, 2 posts)

Researchers are debating how much trust NLA-style mechanistic interpretability outputs deserve, with critics arguing current confidence resembles mere correlation-based readouts.

Episode 2 · Researchers Debate the Safety Value of NLA Interpretability (2026-09-27, 8 posts)

On September 27, AI safety researchers banburismus and thebasepoint clashed across multiple posts over the value of interpretability (interp) research, focusing on how useful NLA (natural language ablation/natural language abstraction) really is for frontier safety.

Confirmed

  • thebasepoint offered a framework for understanding NLA: one of its goals is to extract from activations everything the model could "in principle say," not just what "the assistant would say right now"; he noted that OpenAI's confessor training appears to pursue a similar direction.
  • He offered a rough dichotomy for interpretability research: "how the sausage is made" (the model's internal mechanisms, more elegant) versus "what the sausage is" (the nature of model behavior, empirically more useful for questions about understanding model behavior); he placed NLA roughly ninety percent in the latter camp.
  • He also pointed out NLA's limits: such methods do not uncover internal structure like "addition decomposing into mod-10 components and magnitude components," and are limited in capability for mechanistic discovery.

Why it matters

  • The debate touches on a core divide in interpretability research: dig deep into internal mechanisms, or focus on behavior-level interpretability. thebasepoint's observation is that the latter is empirically more effective while the former is theoretically more elegant, which bears on how safety research resources should be allocated.
  • The link between NLA and OpenAI's confessor training suggests that "getting the model to speak its latent state" is emerging as a notable safety technique, though its mechanistic explanatory power remains contested.