New experiments claim to elicit and observe genuine introspection in LLMs

David Bau's lab, building on the Self-Interpretability protocol, reports eliciting neural signatures of faithful introspection in LLMs; a new COLM paper shows such introspection depends on layer distributions and may enable lie detection.

2026-10-06 ~ 2026-10-06 · 3 related posts