LLMs can access truncated context info; bottleneck is elicitation, not storage
Sauers_ · x · 2026-10-09
Sauers summarizes findings on LLM introspection: models can access information that was in their context but no longer is—e.g., content truncated from context but still present in past chain-of-thought tokens.
- Key takeaway: the limitation is elicitation, not storage—a linear probe can read the information out with high accuracy.
- Mechanism: circuits copy information from past CoT tokens into response tokens, making it accessible on the next turn.
- The same "inhibit introspection" attention head also mediates the Janus / CPU document effect, suggesting a shared underlying mechanism.
More from Research
- Lancet study: patient-facing conversational AI holds up in real urgent care settings — EricTopol · 2026-10-09
- First Workshop on Agent Behavior at COLM 2026 Set for Oct 9 in San Francisco — _Hao_Zhu · 2026-10-09
- Autorubric ships 25-recipe cookbook for rubric design, judge calibration and cost control — deliprao · 2026-10-09
- Autorubric at COLM 2026: a unifying framework for rubric-based LLM evaluation — deliprao · 2026-10-09
- TraceExtract open-sourced: data engine for µ0 world model trained on video with zero action labels — RexDouglass · 2026-10-09
- 500 curated SWE tasks lift Qwen 27B by 11.3 points in 15 GRPO steps — ycombinator · 2026-10-09