New experiments show LLMs can read deleted CoT state — elicitation, not storage, is the bottleneck
Sauers_ · x · 2026-10-09
Sauers released the open-source repo retained-reply-state, testing on Qwen3 models whether the cached KV state of a visible reply carries choices made in hidden thinking, and whether later turns can read it.
- Finding: LLMs have some access to information from tokens that were previously in context
- The bottleneck for this kind of introspection is elicitation, not storage — a linear probe can read it out with high accuracy
- Mechanistically, models use circuits that copy information from past CoT tokens into later representations
- The repo includes full experiment code: attention-head localization, probes, ablations
More from Research
- DNA Typewriter reconstructs mouse embryo lineage: 1.34M profiled cells from zygote to E13.5 — anshulkundaje · 2026-10-09
- Yacine calls for a code reuse benchmark to hill-climb model behavior — yacinelearning · 2026-10-09
- Data companies are becoming research labs, with verifier design the most climbable problem — madhavsinghal_ · 2026-10-09
- Amazon runs Karpathy's AutoResearch at production scale for 12 weeks, finds 5 failure modes — amazon · 2026-10-09
- ReGain: training-free fix restores subject fidelity lost from personalizing on synthetic images — UIUC-CS · 2026-10-09
- OpenAI's claimed Navier–Stokes solution reportedly doesn't match its Lean verification — kyan100 · 2026-10-09