llama.cpp dflash2: Qwen 3.8 27B Inference Speed Up to 3x

Top-Eye-8104 · reddit · 2026-08-20

A user tested the new dflash2 feature in llama.cpp (PR #27342) on an RTX 6000 using Qwen 3.8 27B, comparing four decoding strategies:

Median Results (4 tasks):

DFlash2 delivers an average 3x speedup, though performance varies by task, with one test showing only a 1.5x gain.

Original post →

More from Infra

Infra channel →