New COLM paper: faithful LLMs decide and report with the same layers
dhadfieldmenell · x · 2026-10-06
A new COLM 2026 paper, "Identifying Introspection From the Inside," asks whether LLMs actually know what drives their choices when they report on them — or are just guessing.
Key finding: in the authors' setting, faithful models decide and report using the same layers, while unfaithful ones do not — offering a way to assess introspective faithfulness from the model's internals. The paper will be presented at COLM 2026 Poster Session 1, Imperial Ballroom, poster #66.
More from Research
- 11-Square Packing Optimality Proved and Formalized in Lean With Help From Astra and Claude — ctjlewis · 2026-10-07
- Researcher unveils RSI paradigm: generic disposable agents plus an evolving knowledge base — yisongyue · 2026-10-07
- John Urschel proves Gaussian elimination growth factor is ~n^1/2, settling a Trefethen conjecture — ctjlewis · 2026-10-07
- Where Does Memory Live? RNNs, Transformers, and SSMs Compared Through Working Memory — Pretty_Upstairs9035 · 2026-10-07
- KNOWS Benchmark Jointly Evaluates Agents on Search, Tools and Visual Understanding — anmarasovic · 2026-10-07
- ID-Forcing Keeps Long Video Generation In-Distribution, Enabling Minute-Scale Videos Without Fine-Tuning — _akhaliq · 2026-10-07