Prefix Sliding: 3x Faster LLM Reasoning Without Performance Loss

Muennighoff · x · 2026-08-27

Niklas Muennighoff et al. released "Prefix Sliding for efficient test-time scaling." Addressing the high compute cost of full attention in long-horizon reasoning, the study finds that intermediate tokens lose importance over time. The proposed Prefix Sliding method retains only the essential prefix (instructions/tools) and a sliding window of recent tokens. It achieves 3x speedup without training while maintaining performance, and can scale to over 100k tokens with RL fine-tuning. It outperforms summarization and vanilla sliding window baselines.

Related event: Prefix Sliding triples inference speed with no quality loss(4 posts)→

Original post →

More from Research

Research channel →