Contrastive introspection paper suggests reading LLM neurons to detect lies

davidbau · x · 2026-10-06

David Bau highlights a new paper by diatkinson, dillonplunkett, and himself arguing that the contrastive introspective setup is a powerful experimental platform: it illuminates how large-scale LMs work so well and suggests a path to lie detection. The key lesson: the same AI that gives profound, accurate "I think X" insights can also stochastically parrot and lie about "I think X" — but we may be able to read the neurons to tell the difference. Paper link included.

Related event: New experiments claim to elicit and observe genuine introspection in LLMs(3 posts)→

Original post →

More from Research

Research channel →