New COLM Paper Finds Faithful LLMs Decide and Self-Report With the Same Layers

a_karvonen · x · 2026-10-06

A new COLM paper, "Identifying Introspection From the Inside," asks whether LLMs actually know what drives their decisions or are just guessing when self-reporting. Key finding: faithful models decide and report with the same layers, unfaithful ones don't. LLMs also self-report behaviors from earlier layers much more accurately — a result DeepMind researcher akarvonen calls intuitive and interesting.

Original post →

More from Research

Research channel →