STAIR trains just 12K parameters to turn conversation history into +11.67pp multi-turn reasoning gains
RUC · hf · 2026-10-05
RUC researchers study whether computation from earlier problems helps LLMs solve new ones in the same conversation.
Findings: Retained history can raise or lower later-turn accuracy even within the same domain. Controlled replay experiments isolate problem-history-specific internal state changes, which preserve similar relationships among current problems across histories.
Method STAIR (Stale-Token Attention for Inter-query Reuse):
- Stores keys/values from earlier response generation in a fixed bank;
- Learns to redirect current queries to this bank during prompt processing;
- Base model stays frozen; only 12,288 parameters are trained.
Results: Across three Qwen models and four benchmarks, STAIR improves average later-turn accuracy by up to 11.67 percentage points over the unmodified model with history.
More from Research
- Cohere Labs at COLM: Chain-of-Thought Legibility Is Not Real Interpretability — Cohere_Labs · 2026-10-06
- Gautam Kamath: math community's 'human understanding' push makes AI proofs worth revisiting — thegautamkamath · 2026-10-06
- Everyone uses AI for new math results; Kamath wants AI to simplify old proofs — thegautamkamath · 2026-10-06
- PlurPO: training LLMs to curb social sycophancy that discourages relationship repair — RishiBommasani · 2026-10-06
- First large-scale 3B/8B continuous diffusion LMs match pass@1 and beat pass@k vs masked dLMs — ArashVahdat · 2026-10-06
- AggAgent at COLM: Treats Parallel Agent Trajectories as an Environment for Long-Horizon Tasks — xiye_nlp · 2026-10-06