Qwen3.8-27B: 3x Faster Long-Context Decoding with DFlash2 + XQA

stargate425 · reddit · 2026-08-22

A technical share demonstrates that Qwen3.8-27B achieves a 3x speedup in long-context decoding using DFlash2 + XQA configuration. Tested on an RTX PRO 6000 Max Q at 192K context, the decode speed increased from 18 tok/s to 58 tok/s while remaining lossless in BF16 precision.

Original post →

More from Infra

Infra channel →