How models introspect to recover deleted CoT tokens — and why some do the opposite

Sauers_ · x · 2026-10-09

A research thread explores a striking finding: models can introspect to recover tokens deleted from their past chain-of-thought above chance. The thread digs into the mechanism behind this ability, and why some models use introspection to make it less — not more — likely to say the word they had in mind when asked.

Related event: Study finds LLMs can introspect deleted CoT tokens, with mechanism surprisingly tied to a single attention head(6 posts)→

Original post →

More from Research

Research channel →