How models introspect to recover deleted CoT tokens — and why some do the opposite
Sauers_ · x · 2026-10-09
A research thread explores a striking finding: models can introspect to recover tokens deleted from their past chain-of-thought above chance. The thread digs into the mechanism behind this ability, and why some models use introspection to make it less — not more — likely to say the word they had in mind when asked.
More from Research
- Lancet study: patient-facing conversational AI holds up in real urgent care settings — EricTopol · 2026-10-09
- First Workshop on Agent Behavior at COLM 2026 Set for Oct 9 in San Francisco — _Hao_Zhu · 2026-10-09
- Autorubric ships 25-recipe cookbook for rubric design, judge calibration and cost control — deliprao · 2026-10-09
- Autorubric at COLM 2026: a unifying framework for rubric-based LLM evaluation — deliprao · 2026-10-09
- TraceExtract open-sourced: data engine for µ0 world model trained on video with zero action labels — RexDouglass · 2026-10-09
- 500 curated SWE tasks lift Qwen 27B by 11.3 points in 15 GRPO steps — ycombinator · 2026-10-09