Running DeepSeek V4 Flash on Strix Halo: Vulkan + Speculative Decoding Hits 27 t/s

stereohype · reddit · 2026-08-12

An in-depth benchmark of DeepSeek V4 Flash 0731 on a Flow Z13 (Ryzen AI MAX+ 395, 128GB). Using llama.cpp's Vulkan backend combined with DSpark speculative decoding, it achieves 26.76 t/s sustained decode and 236 t/s prefill.

Cross-platform Comparison

Key Gotchas

Original post →

More from Infra

Infra channel →