New experiments claim to elicit and observe genuine introspection in LLMs
David Bau's lab, building on the Self-Interpretability protocol, reports eliciting neural signatures of faithful introspection in LLMs; a new COLM paper shows such introspection depends on layer distributions and may enable lie detection.
2026-10-06 ~ 2026-10-06 · 3 related posts
- New COLM Paper Finds Faithful LLMs Decide and Self-Report With the Same Layers — a_karvonen · 2026-10-06
- Lab Trains Qwen3-32B to Introspect: Faithful Self-Report Emerges Late, With a Measurable Neural Footprint — davidbau · 2026-10-06
- Contrastive introspection paper suggests reading LLM neurons to detect lies — davidbau · 2026-10-06