WaveFront Decoding Speeds Up Looped Language Models Up to 4.81x Without Retraining
pmttyji · reddit · 2026-10-06
A new paper introduces WaveFront Decoding (WFD), a training-free self-speculative decoding framework for looped language models. It exploits intermediate recurrence outputs as draft predictions and weight sharing to batch token states of different depths into a single recurrent call, forming a diagonal wavefront that drafts and verifies concurrently. On six Spec-Bench task categories, WFD achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, beating draft-then-verify schedules; adding cross-recurrence KV sharing pushes Huginn-3.5B to 4.81x. Code and paper (arXiv:2609.23033) are available on GitHub.
More from Infra
- Google Buys 890 MW of Nuclear Without Building a Single New Reactor — MicahBerkley · 2026-10-06
- Cycle.io Launches DevOps MCP: 3-Node Mongo Replica Set Across 3 Clouds in 15 Minutes — AlexMattoni · 2026-10-06
- Weaviate ships query profiling: a 48ms slow query turned out to be disk reads, not HNSW — CShorten30 · 2026-10-06
- HN Debate: Did Oracle Just Trigger the Implosion of the AI Bubble? — mpweiher · 2026-10-06
- Lambda adopts NVIDIA's AIPerf for model cards showing real-workload inference benchmarks — TheZachMueller · 2026-10-06
- Nvidia nears $6 trillion market value as AI frenzy keeps pushing stocks higher — AryHHAry · 2026-10-06