New COLM Paper Finds Faithful LLMs Decide and Self-Report With the Same Layers
a_karvonen · x · 2026-10-06
A new COLM paper, "Identifying Introspection From the Inside," asks whether LLMs actually know what drives their decisions or are just guessing when self-reporting. Key finding: faithful models decide and report with the same layers, unfaithful ones don't. LLMs also self-report behaviors from earlier layers much more accurately — a result DeepMind researcher akarvonen calls intuitive and interesting.
More from Research
- 0.08 correlation on million-dollar datasets: researchers question virtual cell capital allocation — anshulkundaje · 2026-10-06
- What do LLMs actually mean when they say they're uncertain? A calibration debate — sineadwilliamso · 2026-10-06
- PluginRSI: recursive harness improvement via reusable plugins beats whole-program search — Yaorui Shi · 2026-10-06
- ASCENT: online test-time training lets LLM agents self-distill verified experience into weights — Haodong Lu · 2026-10-06
- Tencent Hunyuan's Prism: dynamic sparse attention speeds up 2K joint video-audio training 2.5x — Tencent-Hunyuan · 2026-10-06
- PaLoRA derives rank-aware pacing law for LoRA continual learning, +4% on 50-task benchmarks — Yuxuan Li · 2026-10-06