WaveFront Decoding Speeds Up Looped Language Models Up to 4.81x Without Retraining

pmttyji · reddit · 2026-10-06

A new paper introduces WaveFront Decoding (WFD), a training-free self-speculative decoding framework for looped language models. It exploits intermediate recurrence outputs as draft predictions and weight sharing to batch token states of different depths into a single recurrent call, forming a diagonal wavefront that drafts and verifies concurrently. On six Spec-Bench task categories, WFD achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, beating draft-then-verify schedules; adding cross-recurrence KV sharing pushes Huginn-3.5B to 4.81x. Code and paper (arXiv:2609.23033) are available on GitHub.

Original post →

More from Infra

Infra channel →