DFlash2 on Qwen3.8 27B hits ~200tk/s for code, requires more VRAM

Hefty_Wolverine_553 · reddit · 2026-08-19

A Reddit user tested Inco AI's DFlash2 speculation decoding with Qwen3.8 27B on an RTX 5090. Results show significant speedups for code generation, peaking at 200 tokens/s and averaging 120 tokens/s per request, compared to previous MTP performance. However, 'thinking' tasks saw a drop to 80-90 tk/s. The trade-off is higher memory usage, forcing a context size reduction from 220k to 160k. Detailed configuration flags were provided.

Original post →

More from Infra

Infra channel →