AMD Strix Halo Runs DeepSeek V4 Flash at 32 tok/s

Developers successfully deployed the 98GB DeepSeek V4 Flash model on a 128GB AMD Strix Halo system. Benchmarks show that utilizing Vulkan backend with speculative decoding achieves local inference speeds exceeding 32 tokens per second.

2026-08-12 ~ 2026-08-12 · 2 related posts