NextLat's latent variable-length speculative decoding speeds inference up to 3.3x

morgymcg · x · 2026-09-02

Because NextLat learns latent dynamics, it unlocks variable-length speculative decoding: drafting a flexible number of future tokens recursively in latent space. Claimed inference speedup of up to 3.3x, faster than MTP-style drafting.

Original post →

More from Research

Research channel →