Strix Halo users ditch official llama.cpp: optimized forks hit ~60 t/s decode vs ~20 t/s
feelspeaceman · reddit · 2026-09-08
The problem
The author estimates 90% of Strix Halo (gfx1151) users run the official llama.cpp, which is not optimized for the chip at all — it struggles to reach 50% of theoretical hardware performance, decoding Qwen 3.8 Flash Next (Q38FN) at only 20+ t/s.
Three alternatives
- peonist-ai/halogen-flash-server: optimized for Strix Halo and Q38FN only, 50 t/s decode and 1200 t/s prefill; described as "Ninfer for Strix Halo"
- myhacsint's experimental llama.cpp fork (strix-halo-qwen4exp): 60 t/s decode and 600 t/s prefill
- halo-box/strix-llama.cpp: 30 t/s decode and 800 t/s prefill; the first continuously updated r/StrixHalo fork with an active Discord, latest commit pushed prefill way up
Each claim links to community user confirmations. The author recommends Q38FN as the go-to model for Strix Halo.
More from Infra
- CPU Shortage Reaches Software Teams as AI Chip Supply Constraints Spread Beyond AI — brada · 2026-09-08
- Big Tech scouts Argentina's Patagonia for new mega data centers — Polymarket · 2026-09-08
- TimesFM 3.0 merges native MLX backend: 641 series/sec on M4 Max, no PyTorch needed — rachittshah · 2026-09-08
- Data center buildout is driving freight demand that the Cass index misses — kernelangus420 · 2026-09-08
- AMD's Halo Station Touts 2.6TB Memory, Tripling NVIDIA's 748GB DGX Station, Shipping 2027 — shashib · 2026-09-08
- DavidSHolz: AI services should let users pay extra to spin up CPU/GPU VMs — DavidSHolz · 2026-09-08