NAVER's DLoop Loops Speculative Decoding Before Verification, Gaining 5-41% Faster Inference Losslessly
naver-ai · hf · 2026-10-08
NAVER AI Lab introduces DLoop, a looped form of speculative decoding that cuts unnecessary target-model forward passes in LLM inference.
- Observation: as draft models get stronger, target models often accept all tokens in a drafting stage, yet verification still follows every stage — wasted computation.
- Method: DLoop performs multiple drafting stages while the draft model stays confident, then verifies all accumulated tokens at once. Loop-aware training keeps the draft model reliable on its own hidden states for unverified tokens.
- Results: across EAGLE-3, DFlash, Domino, DSpark, and multi-token prediction modules, DLoop improves wall-clock speedup by 5-41% with lossless decoding. Code to be released on GitHub.
More from Infra
- OpenRSI agents land 3 optimizations merged upstream into SGLang-Omni — yuz9yuz · 2026-10-08
- Photonics engineer asks why waveguide facet angles stop short with a perpendicular portion on transmitter chips — jwt0625 · 2026-10-08
- China's electricity glut turns data centers into a solution, as 14nm chips get pressed into service — teortaxesTex · 2026-10-08
- Transformer lead times balloon from 500 to 1,120 days, YC partner calls it a startup opportunity — ycombinator · 2026-10-08
- MIT's Christina Delimitrou uses AI to cut data center energy waste and downtime — nordicinst · 2026-10-08
- FT kicks off three-part series on China's breakneck AI infrastructure build-out, from Ulanqab to Shaoguan — zijing_wu · 2026-10-08