NextLatent teaches transformers to predict their own latent states, enabling 3.3x faster inference

burny_tech · x · 2026-09-05

Jayden Teoh and collaborators present Next-Latent Prediction (NextLat), a self-supervised method addressing next-token prediction's myopia: transformers learn to predict their own next latent state, forming compact world models for reasoning and planning. The training pressure pushes models toward belief-state representations, and the approach unlocks up to 3.3x faster inference via self-speculative decoding. Yacine says he'll interview the first author next week.

Original post →

More from Research

Research channel →