AMD Strix Halo Runs DeepSeek V4 Flash at 32 tok/s
Developers successfully deployed the 98GB DeepSeek V4 Flash model on a 128GB AMD Strix Halo system. Benchmarks show that utilizing Vulkan backend with speculative decoding achieves local inference speeds exceeding 32 tokens per second.
2026-08-12 ~ 2026-08-12 · 2 related posts
- Running DeepSeek V4 Flash on Strix Halo: Vulkan + Speculative Decoding Hits 27 t/s — stereohype · 2026-08-12
- DeepSeek V4 Flash Hits 32.7 tok/s on AMD Strix Halo via 98GB GGUF — tensorqt · 2026-08-12