SpecFold exploits multi-branch redundancy to speed diffusion LLM decoding up to 1.99x
GeorgiaTech · hf · 2026-10-10
GeorgiaTech researchers propose SpecFold, an algorithm-system co-design that accelerates multi-branch speculative decoding for diffusion LLMs (DLLMs).
- Key insight: prior acceleration exploits temporal redundancy across denoising steps; SpecFold targets redundancy within speculative verification—draft branches inherit most tokens from parents, leaving hidden states highly similar across branches.
- Method: token-level residual gating with folded attention and FFN to reuse parent computation, plus a Triton kernel for sparse multi-branch execution.
- Results: across 2 DLLM families, 5 models, and 5 benchmarks, up to 1.64x throughput over Spiffy and 1.99x over vanilla decoding with comparable task performance; orthogonal to temporal caching.
More from Infra
- Dresden fab likely targets 7nm without EUV via immersion multi-patterning — pstAsiatech · 2026-10-10
- Samsung open-sources LittleBit: extreme quantization fits a 13B model in under 1GB — udmrzn · 2026-10-10
- Analyst says agentic CPU research lines up with NVIDIA Vera benchmark, flags scale-up vs scale-out question — BenBajarin · 2026-10-10
- Container cuts agent runtime P95 latency to 731ms, down ~180ms in two weeks — ritakozlov · 2026-10-10
- Amazon drops data center NDAs as community backlash spurs hundreds of moratoriums — TechCrunch AI · 2026-10-10
- Amazon drops data center NDAs, and AI agents want your credit card — TechCrunch AI · 2026-10-10