Interpretability Thread: What NLA Could Extract From Model Activations, and OpenAI's Confessor Training

thebasepoint · x · 2026-09-27

In an interpretability discussion, the author proposes one framing of NLA: extracting everything the model could in principle say from an activation, independent of what 'the assistant' would say. He notes OpenAI's confessor training looked like an attempt to enable exactly this, but nothing beyond a prototype has surfaced since January.

Related event: Researchers Debate the Value of NLA and Two Routes of Interpretability(5 posts)→

Original post →

More from Research

Research channel →