Training-free WaveFront Decoding speeds up looped LMs by up to 4.81x
SNU-VLSI · hf · 2026-09-29
WaveFront Decoding is a training-free self-speculative decoding scheme for looped LMs that batches drafting and verification in the same recurrent calls. It achieves 2.42x speedup on Ouro-2.6B and up to 4.81x on Huginn-3.5B with cross-recurrence KV sharing.
More from Infra
- Qdrant unveils Constella research preview: swap query embedding models without re-embedding your docs — qdrant_engine · 2026-09-29
- 124M model with a 65B embedding sparks the AFED disaggregation joke — YouJiacheng · 2026-09-29
- Oracle's 30-year spread widens to +180bp as Project Jupiter power woes trigger force majeure — julsimon · 2026-09-29
- Bain: AI needs $6T annual revenue by 2031 to justify data-center spending, $4.2T gap remains — rohanpaul_ai · 2026-09-29
- Cost math: Meta Muse would run $51 per user, making 'free for 4B users' a steep climb — bookwormengr · 2026-09-29
- PReCache: Training-Free KV Cache Sharing Gives Multi-LoRA Agents up to 3.1x TTFT Speedup — SNU-VLSI · 2026-09-29