llama.cpp's New DFlash2 Boosts Qwen Inference Speed Up to 3x

Community tests show llama.cpp's new DFlash2 feature speeds up Qwen 3.8 27B inference by up to 3x on RTX 6000, though one RTX 5090 test found it 20% faster than MTP at the cost of a 38% smaller context.

2026-08-20 ~ 2026-08-21 · 2 related posts

Full story(14 episodes)→