Models introspect to recover tokens deleted from past CoT — but the mechanism is a mystery

Sauers_ · x · 2026-10-09

Opening post of Sauers' introspection research thread: experiments show models can introspect to recover tokens deleted from their past chain-of-thought at above-chance rates. Two open questions remain — how is this possible mechanistically, and why do some models use their introspection to make it less (not more) likely to answer the word they thought of when asked? Later posts in the thread unpack the experimental design and the "hidden animal" behavior.

Related event: Study finds LLMs can introspect deleted CoT tokens, with mechanism surprisingly tied to a single attention head(6 posts)→

Original post →

More from Research

Research channel →