Strix Halo users ditch official llama.cpp: optimized forks hit ~60 t/s decode vs ~20 t/s

feelspeaceman · reddit · 2026-09-08

The problem

The author estimates 90% of Strix Halo (gfx1151) users run the official llama.cpp, which is not optimized for the chip at all — it struggles to reach 50% of theoretical hardware performance, decoding Qwen 3.8 Flash Next (Q38FN) at only 20+ t/s.

Three alternatives

Each claim links to community user confirmations. The author recommends Q38FN as the go-to model for Strix Halo.

Original post →

More from Infra

Infra channel →