gufo inference doubles prefill speed vs llama.cpp forks for Qwen 3.8 Flash Next on Strix Halo

fallingdowndizzyvr · reddit · 2026-09-28

A Reddit user (not affiliated) recommends the open-source gufo inference framework for running Qwen 3.8 Flash Next on AMD Strix Halo, saying it's much faster than llama.cpp at high context. Real-workload chat numbers: 6204 chunks in 119s, encode 1239 tok/s, decode 57 tok/s with MTP on — prefill (PP) roughly 2x the fastest Strix Halo-specific llama.cpp fork and far beyond mainline. gufo supports other models too, though the list is still short.

Original post →

More from Infra

Infra channel →