Qwen3.8-27B: 3x Faster Long-Context Decoding with DFlash2 + XQA
stargate425 · reddit · 2026-08-22
A technical share demonstrates that Qwen3.8-27B achieves a 3x speedup in long-context decoding using DFlash2 + XQA configuration. Tested on an RTX PRO 6000 Max Q at 192K context, the decode speed increased from 18 tok/s to 58 tok/s while remaining lossless in BF16 precision.
More from Infra
- Proposal to convert offshore oil rigs into data centers — beffjezos · 2026-08-22
- OpenAI DNS Records Hint at Parallel Banking and Hardware Infrastructure — imjustnewatai · 2026-08-22
- AI energy crisis: ChatGPT queries use 10x energy of Google search — ingliguori · 2026-08-22
- Google Cloud launches Global Front End for cross-cloud networking — rseroter · 2026-08-22
- Opinion: 'Ban Data Centers' is a Luxury Belief That Would Disastrously Impact Economy — robleclerc · 2026-08-22
- Path to 100x AI Efficiency? Needs Architecture Shift, Says VC — prateekj · 2026-08-22