gufo inference doubles prefill speed vs llama.cpp forks for Qwen 3.8 Flash Next on Strix Halo
fallingdowndizzyvr · reddit · 2026-09-28
A Reddit user (not affiliated) recommends the open-source gufo inference framework for running Qwen 3.8 Flash Next on AMD Strix Halo, saying it's much faster than llama.cpp at high context. Real-workload chat numbers: 6204 chunks in 119s, encode 1239 tok/s, decode 57 tok/s with MTP on — prefill (PP) roughly 2x the fastest Strix Halo-specific llama.cpp fork and far beyond mainline. gufo supports other models too, though the list is still short.
More from Infra
- US AI datacenter spending as share of GDP now exceeds historic railroad and highway buildouts combined — rvp · 2026-09-28
- FailureAtlas: most severe LLM gateway failures return HTTP 200 and silently corrupt your app — its_vayishu · 2026-09-28
- Used RTX 3090 prices creep toward $1,500 on eBay amid GPU shortage — sleight42 · 2026-09-28
- Scraping p50 stabilized at 2s: keep your app and databases colocated — DanielLockyer · 2026-09-28
- PKU open-sources RayOrch, lineage-aware data-prep engine with up to 15.14x speedup — PekingUniversity · 2026-09-28
- Why credit quality matters most in compute: clusters are opaque, people are trackable — AccBalanced · 2026-09-28