DeepSeek-V4-Flash Ported to Run on AMD Strix Halo APU

Fit-Produce420 · reddit · 2026-07-31

A developer attempted to run a quantized DeepSeek-V4-Flash-GGUF model on AMD's latest Strix Halo hardware.

According to the feedback, the model can support a context length of approximately 64K on a single Strix Halo, achieving an inference speed of about 35 tokens/s. While not exceptionally fast, this local deployment setup is considered practically usable if the model's capabilities hold up.

Original post →

More from Infra

Infra channel →