Qwen3.8-27B hits 31t/s on AMD Strix Halo

stereohype · reddit · 2026-08-20

On a device with Ryzen AI Max+ 395 and Radeon 8060S, the author achieved 31.4 t/s decode speed for Qwen3.8-27B (Q5KXL) at 80W using the DFlash2 speculative decoding and Vulkan backend, with 300 t/s prefill. Benchmarks show DFlash2 is 40% faster than the built-in MTP, and surprisingly, Q5 quantization is faster than Q4 in this setup due to higher draft acceptance rates. Full startup commands and optimization tips are provided.

Original post →

More from Infra

Infra channel →