Qwen 27B Now Runs on AMD NPUs via FastFlowLM, at a Slow 1 tps

TuskNaPrezydenta2020 · reddit · 2026-10-01

FastFlowLM v1.0.7 now runs the 27B Qwen model locally on AMD NPUs—though the poster jokes the decode speed is just 1 token per second. Open source under ROCm/FastFlowLM on GitHub.

Original post →

More from Infra

Infra channel →