Microsoft & Cornell Paper: Free Pause Tokens Boost Prediction With ~1.14x Training-Only Overhead
dair_ai · x · 2026-09-06
dairai highlights a Microsoft/Cornell paper on pause tokens: they buy the model extra compute per next-token prediction, paying only with a sequence position. Free pause tokens carry the same compute in a parallel prediction stream over a weight-shared backbone. At inference they add nothing to context length, leave the KV cache unchanged, and cost essentially no latency; a 1B model gains 2-3 centinats on next-token prediction. The cost moves to training — roughly 1.14x overhead versus an optimized pretraining pipeline. A notable alternative if thinking-token context costs are unaffordable.
More from Infra
- GPU programming diarist digs into physical memory systems after studying pipeline parallelism — blaizedsouza · 2026-09-06
- Solo Dev Spent Six Months on an On-Device Android Agent Built as an Editable Graph, Not a Prompt — alexeyw · 2026-09-06
- Time-encoded analog photonic interposer links vision sensors to analog accelerators, 2.04x EDP gain — jwt0625 · 2026-09-06
- Conviva Swapped mmap for io_uring in Its Rust Query Engine — It Got Slower — blaizedsouza · 2026-09-06
- DLSS5 video upscaling is lightning fast—could multipass x2 loops replace AI upscalers? — More-Ad5919 · 2026-09-06
- That PyTorch matmul Precision Warning Is Worth Reading After All — generativist · 2026-09-06