New papers show LLM introspection splits into distinct detection and identification mechanisms

gsarti_ · x · 2026-10-10

New work on LLM introspection: a COLM paper ('Identifying Introspection From the Inside') and a replication study (Lederman & Mahowald) show injected-thought detection splits into distinct anomaly-detection and identification mechanisms in different layers; introspection is 'content-agnostic' — models detect anomalies but confabulate injected concepts like 'apple'; and only 'verbalizable' representations, tied to a global-workspace account, support faithful introspective reports.

Original post →

More from Research

Research channel →