Models introspect to recover tokens deleted from past CoT — but the mechanism is a mystery
Sauers_ · x · 2026-10-09
Opening post of Sauers' introspection research thread: experiments show models can introspect to recover tokens deleted from their past chain-of-thought at above-chance rates. Two open questions remain — how is this possible mechanistically, and why do some models use their introspection to make it less (not more) likely to answer the word they thought of when asked? Later posts in the thread unpack the experimental design and the "hidden animal" behavior.
More from Research
- China Telecom and MemTensor unveil HaluMem, first operation-level benchmark for agent memory hallucinations — jiqizhixin · 2026-10-09
- BAAI's AREX research agent checks answers requirement-by-requirement, hits 82.5% BrowseComp — DeepLearningAI · 2026-10-09
- New piece lays out how to build RL environments aimed at superintelligence — JenniferHli · 2026-10-09
- Quantum Counterfactuals: Quantum RNGs as an Exploration Source for RL — jessi_cata · 2026-10-09
- OpenAI theorem drop collides with researchers' work: stronger bounds but 'unreadable' proof — guyvdb · 2026-10-09
- AI-written science floods preprint servers; researchers propose decision language models as filter — lpachter · 2026-10-09