Reasoning trace placement swings long-context accuracy by 50 points, Trace-as-State technique shows
dair_ai · x · 2026-09-04
A practical technique for reasoning models: place the reasoning trace before the long-context block on a fresh pass instead of appending it after. Because Transformers process causally, earlier-derived state can guide rereading. On GraphWalks Parents, DeepSeek V4 Pro Preview jumps from 29.2% (baseline) and 43.0% (trace appended) to 81.8% with Trace as State; GLM-5.2 improves from 66.4% / 83.2% similarly. The author notes that providing conditions first can require exponentially less memory in the worst case for causal processors. A near-free context-ordering change with huge accuracy payoff.
More from Models
- OpenAI benchmarking against Claude models again, and the AI world notices — hrishioa · 2026-09-04
- VulcanBench-SWE v4 Raises Timeout to 10 Hours to Benchmark New Coding Models Cleanly — ChrisUniverse · 2026-09-04
- Models show significant, continued progress in computational bio and statistical reasoning — anshulkundaje · 2026-09-04
- KOL slams OpenAI's 'reckless' Astra release as a move to bully Anthropic — scaling01 · 2026-09-04
- Claimed GPT-6 'Astra' saturates ARC-AGI-3 as Brockman says 'welcome to the AGI era' — ShafeDogg · 2026-09-04
- Claude Defends User's Bad Architecture Decisions in the Name of 'Honesty' — repligate · 2026-09-04