Optimized SPS yields 'free pause tokens' that boost transformer inference at 10-20% overhead, down from ~700%
JohnCLangford · x · 2026-09-08
John Langford and collaborators (including @nthngdy and Yoav Artzi) ran State Prediction Separation through an optimization process, finding it works at a higher operating point and yielding 'free pause tokens' that let a transformer make higher-quality inference at near-zero cost.
Training-time overhead drops from 700% in the original SPS to an amortized 10-20% depending on design choices — an improvement on basic transformers the authors say may be widely useful.
More from Research
- AI gender-violence detection study misflags 46% of non-survivors in test group — kmcolo · 2026-09-09
- A Millennium Prize Problem reportedly solved — with a spicy human backstory — mmbronstein · 2026-09-09
- Nature Reviews Cancer at 25: researchers weigh agentic AI and human-AI co-science in oncology — marinkazitnik · 2026-09-09
- Causal foundation models estimate causal effects in-context, no fine-tuning needed — Layer6 · 2026-09-09
- Conformal Relevance framework automates conformal score design via in-context ensembles — Layer6 · 2026-09-09
- Omnii, a language model pretrained on DNA, designs personalized mRNA cancer vaccines — exnx · 2026-09-09