llama.cpp dflash2: Qwen 3.8 27B Inference Speed Up to 3x
Top-Eye-8104 · reddit · 2026-08-20
A user tested the new dflash2 feature in llama.cpp (PR #27342) on an RTX 6000 using Qwen 3.8 27B, comparing four decoding strategies:
Median Results (4 tasks):
- Baseline: 47.4 tok/s
- MTP: 114.7 tok/s
- DFlash: 99.3 tok/s
- DFlash2: 140.6 tok/s
DFlash2 delivers an average 3x speedup, though performance varies by task, with one test showing only a 1.5x gain.
More from Infra
- Ramp launches Router, an OpenRouter competitor, after 3 years of internal use — himanshustwts · 2026-08-20
- Rise of S3-native products as a new developer trend — DanielLockyer · 2026-08-20
- Terminal tool cmux reported for high memory usage — DanielLockyer · 2026-08-20
- Loudoun County Gets Rich From Data Centers, Residents Question the Cost — AndyMasley · 2026-08-20
- Developer Hits One Billion Tokens in a Single Week — jh3yy · 2026-08-20
- Local data centers face NIMBY backlash as 'dystopian betrayal' — ai · 2026-08-20