New papers show LLM introspection splits into distinct detection and identification mechanisms
gsarti_ · x · 2026-10-10
New work on LLM introspection: a COLM paper ('Identifying Introspection From the Inside') and a replication study (Lederman & Mahowald) show injected-thought detection splits into distinct anomaly-detection and identification mechanisms in different layers; introspection is 'content-agnostic' — models detect anomalies but confabulate injected concepts like 'apple'; and only 'verbalizable' representations, tied to a global-workspace account, support faithful introspective reports.
More from Research
- Toronto surgeons train AI to flag safe incision zones in real time during surgery — EricTopol · 2026-10-11
- DeepSeek-V4.1-Flash: bycloud breaks down DeepSeek's boldest architecture overhaul yet — bycloud · 2026-10-11
- Mathematician digests OpenAI's number theory results; Hodge papers pulled over sign error — lpachter · 2026-10-11
- SpIDER paper boosts code retrieval for coding agents via semantic search plus code graphs — mangahomanga · 2026-10-11
- CMU professor builds detailed 3D dragon from 27KB of code via Astra — 141_1337 · 2026-10-11
- Claude surfaces hidden planetary system 158 light-years away from public telescope data — DavidmComfort · 2026-10-11