DFlash 2 speculative decoding hits 2.26x on real coding, 4.68x stacked with n-gram — 3-day llama.cpp benchmark
FantasticNature7590 · reddit · 2026-08-23
A 3-day llama.cpp benchmark (PR #27342) pits Inco AI's DFlash 2 against MTP and n-gram drafters on Qwen 3.8 27B (single RTX PRO 6000, concurrency 1):
- DFlash 2 alone: 67.97→153.91 tok/s (2.26x) on 100 real LiveCodeBench problems, ITL 14.27→6.02ms, natural stops, +2.7GB VRAM. Beats DFlash 1's 2.00x at matched draft width for roughly half the VRAM.
- DFlash 2 + one n-gram table (k4v): 4.68x on the build phase of an 18-turn coding session (65.1→304.9 tok/s); adding the second table dropped it to 3.77x—flipping DFlash 1's July result.
- Parameter traps: recommended --spec-draft-n-max 7 is past peak (5 gives 11% more on 8K coding prompts) and silently clamped; --spec-draft-p-min is never read in the DFlash 2 code path; the +52% synthetic gain is harness degeneration, and n-gram is -30% on prose.
- Full reproducible config with quant versions, cooling and anti-throttling measures; the author honestly flags the cross-generation rows aren't a controlled A/B.
More from Infra
- 25% chance orbital data centers launch by end of next year — Polymarket · 2026-08-23
- VC proposes network of giant data centers along US-Mexico border — Polymarket · 2026-08-23
- Experiment proposed: Local Qwen model on Mac vs $10k cloud security scan — natesiggard · 2026-08-23
- Dual R9700 vLLM Setup Halves Speed with Multiple Instances — Certain_Series6810 · 2026-08-23
- Best Local LLMs for 12GB VRAM: Alternatives to GLM 4.7 Flash? — OrangeThink5911 · 2026-08-23
- Polymarket: 12% chance AI bubble bursts by year-end — Polymarket · 2026-08-23