Optimized SPS yields 'free pause tokens' that boost transformer inference at 10-20% overhead, down from ~700%

JohnCLangford · x · 2026-09-08

John Langford and collaborators (including @nthngdy and Yoav Artzi) ran State Prediction Separation through an optimization process, finding it works at a higher operating point and yielding 'free pause tokens' that let a transformer make higher-quality inference at near-zero cost.

Training-time overhead drops from 700% in the original SPS to an amortized 10-20% depending on design choices — an improvement on basic transformers the authors say may be widely useful.

Related event: Microsoft researchers find optimized "free pause tokens" boost LLM reasoning at near-zero cost(2 posts)→

Original post →

More from Research

Research channel →