Strix Halo Plus R9700 Hybrid Hits 60+ tok/s, Beating DGX Spark on Local LLM Inference
darklordfireape · reddit · 2026-10-01
After months of iteration, the MIT-licensed llama-halo-hybrid project lets you pair an R9700 GPU (via PCIe riser, OCuLink, or Thunderbolt dock) with a Strix Halo APU: dense layers, KV cache and some layers go on the GPU, the rest stays on the APU. It's a modified llama.cpp requiring no custom quants, tuned for Qwen and GLM families. Running Qwen-3.8-flash-next at Q4/Q4KXL it breaks 60+ tok/s decode and 2000+ tok/s prefill with full 256k context — beating the pricier DGX Spark, with the Swift-1.5 variant doing even better. Author suggests gufo for pure Strix Halo users and invites tuning requests for other sidecar GPUs.
More from Infra
- DRAM supply shows no line of sight to catching up with demand, says analyst — BenBajarin · 2026-10-01
- Micron: humanoid robots may need memory comparable to autonomous vehicles, driving demand by 2030 — McDonaghMatthew · 2026-10-01
- antirez praises 192GB Framework Desktop, plugs external GPU for heterogeneous inference — antirez · 2026-10-01
- RTX 5090 + Intel Arc B70 for local LLMs: halved throughput vs bigger context — indiealexh · 2026-10-01
- Google explores space data centers: keeping AI chips from overheating in orbit — GraceToSentience · 2026-10-01
- Magnitude inference engine hits #1 on HN, claims up to 2x faster local open-model runs than llama.cpp — nickbaumann_ · 2026-10-01