STAIR trains just 12K parameters to turn conversation history into +11.67pp multi-turn reasoning gains

RUC · hf · 2026-10-05

RUC researchers study whether computation from earlier problems helps LLMs solve new ones in the same conversation.

Findings: Retained history can raise or lower later-turn accuracy even within the same domain. Controlled replay experiments isolate problem-history-specific internal state changes, which preserve similar relationships among current problems across histories.

Method STAIR (Stale-Token Attention for Inter-query Reuse):

Results: Across three Qwen models and four benchmarks, STAIR improves average later-turn accuracy by up to 11.67 percentage points over the unmodified model with history.

Original post →

More from Research

Research channel →