Interpretability Thread: What NLA Could Extract From Model Activations, and OpenAI's Confessor Training
thebasepoint · x · 2026-09-27
In an interpretability discussion, the author proposes one framing of NLA: extracting everything the model could in principle say from an activation, independent of what 'the assistant' would say. He notes OpenAI's confessor training looked like an attempt to enable exactly this, but nothing beyond a prototype has surfaced since January.
Related event: Researchers Debate the Value of NLA and Two Routes of Interpretability(5 posts)→
More from Research
- MIT economist challenges Aaru's human-simulation benchmarks: opaque method, no baseline, data leakage risks — soumitrashukla9 · 2026-09-27
- DeMiAn: dense language annotations boost robot policy learning, cut compute 62% — rajammanabrolu · 2026-09-27
- UC Berkeley opens tenure-track faculty position in AI for Biology — anshulkundaje · 2026-09-27
- Single neuron sufficient to bypass safety alignment in LLMs, paper finds — amplifiedamp · 2026-09-27
- C5R's SciUniverse benchmark exposes AI failures at the lab bench — VraserX · 2026-09-27
- Interpretability researcher: sandbagging signals from probes would block model deployment — thebasepoint · 2026-09-27