DFlash 2 speculative decoding hits 2.26x on real coding, 4.68x stacked with n-gram — 3-day llama.cpp benchmark

FantasticNature7590 · reddit · 2026-08-23

A 3-day llama.cpp benchmark (PR #27342) pits Inco AI's DFlash 2 against MTP and n-gram drafters on Qwen 3.8 27B (single RTX PRO 6000, concurrency 1):

Original post →

More from Infra

Infra channel →