Study finds LLMs can introspect deleted CoT tokens, with mechanism surprisingly tied to a single attention head

Mechanistic interpretability researcher Sauers posted a thread on October 9 and open-sourced the experiment repo retained-reply-state, systematically studying LLM introspection: models can recover, at above-chance rates, tokens deleted from their prior chains of thought (CoT) by "introspecting"—showing the information still lingers in the cached state of visible replies (KV cache), which later turns can read. This finding matters because it directly touches on whether a model's "inner state" can be read by itself and by external auditors.

Confirmed

Not yet confirmed

Why it matters

2026-10-09 ~ 2026-10-09 · 6 related posts

Primary sources