DLoop: looped speculative decoding cuts target-model passes, boosting speedup 5-41% across EAGLE-3 and more

pmttyji · reddit · 2026-10-09

NAVER AI's DLoop paper observes that strong draft models often get all tokens accepted, yet verification still runs after every drafting stage. DLoop adaptively performs multiple drafting stages while the draft model stays confident, then verifies accumulated tokens in one pass; loop-aware training keeps drafting reliable on its own unverified hidden states. Across EAGLE-3, DFlash, Domino, DSpark and multi-token prediction modules, wall-clock speedup improves 5-41% with lossless decoding. Code is coming soon at github.com/naver-ai/DLoop.

Related event: NAVER's DLoop Speeds Up Speculative Decoding by up to 41%(2 posts)→

Original post →

More from Infra

Infra channel →