Strix Halo Plus R9700 Hybrid Hits 60+ tok/s, Beating DGX Spark on Local LLM Inference

darklordfireape · reddit · 2026-10-01

After months of iteration, the MIT-licensed llama-halo-hybrid project lets you pair an R9700 GPU (via PCIe riser, OCuLink, or Thunderbolt dock) with a Strix Halo APU: dense layers, KV cache and some layers go on the GPU, the rest stays on the APU. It's a modified llama.cpp requiring no custom quants, tuned for Qwen and GLM families. Running Qwen-3.8-flash-next at Q4/Q4KXL it breaks 60+ tok/s decode and 2000+ tok/s prefill with full 256k context — beating the pricier DGX Spark, with the Swift-1.5 variant doing even better. Author suggests gufo for pure Strix Halo users and invites tuning requests for other sidecar GPUs.

Original post →

More from Infra

Infra channel →