SpecFold exploits multi-branch redundancy to speed diffusion LLM decoding up to 1.99x

GeorgiaTech · hf · 2026-10-10

GeorgiaTech researchers propose SpecFold, an algorithm-system co-design that accelerates multi-branch speculative decoding for diffusion LLMs (DLLMs).

Original post →

More from Infra

Infra channel →