Reusing reasoning traces as pre-context state lifts long-context accuracy in 26 of 27 tests
burny_tech · x · 2026-09-04
The paper "Trace as State: Reasoning Traces as Conditional States for Long-Context Transformers" tackles a limitation: task state discovered late in the context can't influence how earlier tokens were encoded.
The fix: reuse the model's own reasoning trace as task state, place it before the context, and run a fresh pass so the model re-reads with that state already available. This beats appending the same trace after the context in 26/27 settings, with DeepSeek V4 Pro jumping from 43.0% to 81.8% EM on GraphWalks Parents.
More from Research
- IP-Adapter at 0.6 overrides prompts; lower it and 4-view character consistency breaks — God_Speedmyboy · 2026-09-04
- 7M-Parameter Tiny Recursion Model Hits 45% on ARC-AGI-1 — CatAstro_Piyush · 2026-09-04
- New Paper Tunes Training-Time MSA Depth Distribution to Beat AlphaFold2/3 — MoAlQuraishi · 2026-09-04
- KV cache doesn't enable information passing, it just saves compute — scaling01 · 2026-09-04
- YC-affiliated team offers $500k and free YAM arms for dexterous manipulation benchmarks — ycombinator · 2026-09-04
- MITRA team lands two WMT2026 papers: largest Sanskrit-English corpus and 9B translation model — SebastianNehrd2 · 2026-09-04