Prefix Sliding: Discard Intermediate Reasoning Tokens to Cut Test-Time Compute, Training-Free

burkov · x · 2026-09-08

A paper introduces Prefix Sliding, addressing test-time scaling's costs: full attention keeps every prior token, so compute grows linearly with length, plus distraction by irrelevant tokens, repetitive loops, and lost information. The question: can most intermediate reasoning tokens be discarded without harming results?

Original post →

More from Models

Models channel →