Qwen3.8-27B hits 31t/s on AMD Strix Halo
stereohype · reddit · 2026-08-20
On a device with Ryzen AI Max+ 395 and Radeon 8060S, the author achieved 31.4 t/s decode speed for Qwen3.8-27B (Q5KXL) at 80W using the DFlash2 speculative decoding and Vulkan backend, with 300 t/s prefill. Benchmarks show DFlash2 is 40% faster than the built-in MTP, and surprisingly, Q5 quantization is faster than Q4 in this setup due to higher draft acceptance rates. Full startup commands and optimization tips are provided.
More from Infra
- Qwen3.8-27B Test: Lower KV Cache Quantization Impacts Reasoning Quality — fbms2 · 2026-08-20
- Trading speed for cost on DGX Spark, RTX 5090 suggested for speed boost — daniel_mac8 · 2026-08-20
- 1.5-Year Delay Cuts AI Data Center Value by 8.9%, Speed Key: Report — Beth_Kindig · 2026-08-20
- Clarification: OpenAI's 20% compute claim refers to monitoring overhead, not total capacity — sjgadler · 2026-08-20
- SkyPilot, VAST Data, and Partners Host AI Infra Meetup — skypilot_org · 2026-08-20
- PolymathicAI Releases The Well: A 15TB Collection of Physics Simulations — tom_doerr · 2026-08-20