Strix Halo + dGPU real-world test: low-context benchmarks oversell the speedup

Hrethric · reddit · 2026-08-25

The author attached an R9700 dGPU to a Strix Halo machine via Oculink. Running Unsloth quants of MiniMax M-2.7, Q4KXL on the dGPU was slower than IQ4XS on the APU alone; after hours of tuning, IQ4XS matched APU token generation with a 2-3x prompt-processing boost and a modest context bump.

Many published dGPU speedups were run at low context, he notes. There is a real tradeoff: offload more layers to the dGPU at the cost of context size/quality, or keep a higher-quality context with fewer dGPU layers. Full Docker command and models.ini settings (tensor-split, cache-type, batch sizes) are included.

Original post →

More from Infra

Infra channel →