New COLM paper: faithful LLMs decide and report with the same layers

dhadfieldmenell · x · 2026-10-06

A new COLM 2026 paper, "Identifying Introspection From the Inside," asks whether LLMs actually know what drives their choices when they report on them — or are just guessing.

Key finding: in the authors' setting, faithful models decide and report using the same layers, while unfaithful ones do not — offering a way to assess introspective faithfulness from the model's internals. The paper will be presented at COLM 2026 Poster Session 1, Imperial Ballroom, poster #66.

Related event: Bau Lab Induces and Observes Faithful Introspection in LLMs for the First Time(6 posts)→

Original post →

More from Research

Research channel →