Single Strix Halo Beats Six-GPU Rig Running 176B Qwen MoE at Long Context
fallingdowndizzyvr · reddit · 2026-10-01
A Reddit user benchmarked AMD's Strix Halo APU against a pile of consumer GPUs (2x RTX 5070 Ti, 2x RX 7900 XTX, 2x RTX 5060 Ti 16GB) running Qwen QFN (176B A3B MoE, Q4 quantization, 104GB).
The counterintuitive result: at 160K context, the single Strix Halo wins decisively:
- Six-GPU rig: PP 215.73 / TG 16.29 t/s
- Strix Halo (Gufo backend): PP 1227.12 / TG 22.04 t/s
- Strix Halo (Halo Box fork): PP 587.99 / TG 21.24 t/s
- Strix Halo (llama.cpp mainline 0.4.1): PP 113.48 / TG 7.06 t/s
The consumer GPUs lack VRAM, so pipeline parallelism and cross-card communication crush long-context throughput, while Strix Halo's unified memory shines. Backend choice also matters enormously: the Gufo-optimized build delivers 3x+ the TG speed of llama.cpp mainline.
More from Infra
- Delip Rao: Most big-budget GPU training runs are run sub-optimally — deliprao · 2026-10-01
- Tencent Leases 100,000 Chips From Oracle, Ex-OpenAI Exec Calls It Insane — Miles_Brundage · 2026-10-01
- Silicon microring modulators push past 200Gb/s per lane to cut AI optical I/O power — jwt0625 · 2026-10-01
- WUSH-KV: Data-Adaptive Transforms for 2-bit KV-Cache Quantization Integrated into SGLang — ISTA-DASLab · 2026-10-01
- Linewise's video agent hits 31x GPU throughput at 1/15 cost on Inco inference infra — songhan_mit · 2026-10-01
- Memory stocks rally overnight in Asia as AI demand frenzy reignites — firstadopter · 2026-10-01