DFlash Merges into llama.cpp for 4.44x Faster Local Inference

FantasticNature7590 · reddit · 2026-07-08

Developers tested the newly merged DFlash speculative decoding (PR #22105) in llama.cpp on an RTX 6000 PRO. Running Qwen 3.6 27B at a 36K context, it proved to be 4.44 times faster than the previous best MTP. Developed by z-lab, DFlash uses a block diffusion drafter that fills a 15-token block at once instead of generating tokens individually. It can be deployed as a Llama server via a one-click Docker setup.

Related event: llama.cpp Integrates DFlash Speculative Decoding for Major Local Inference Speedup(5 posts)→

Original post →

More from Infra

Infra channel →