Qwen3.8-27B achieves 2x decode speedup with DFlash2 on 256k context

maddie-lovelace · reddit · 2026-08-21

Tests on Qwen3.8-27B-UD-Q4KXL with DFlash2 on a single RTX 5090 show roughly 2x decode speedup (40 to 75 tps) at 256k context with only a 15% prefill slowdown. The post includes specific reproduction commands and configuration notes.

Original post →

More from coding & agent

coding & agent channel →