Microsoft's John Langford finds optimized 'free pause tokens' boost transformer inference near-free

JohnCLangford · x · 2026-09-08

John Langford shares a fun paper: his team ran State Prediction Separation through an optimization process, finding it works at higher op and yields 'free pause tokens' that let a transformer make higher-quality inference at near-zero extra cost.

A lightweight approach to letting models think longer without a meaningful inference cost increase.

Related event: Microsoft researchers find optimized "free pause tokens" boost LLM reasoning at near-zero cost(2 posts)→

Original post →

More from Research

Research channel →