Strix Halo + dGPU real-world test: low-context benchmarks oversell the speedup
Hrethric · reddit · 2026-08-25
The author attached an R9700 dGPU to a Strix Halo machine via Oculink. Running Unsloth quants of MiniMax M-2.7, Q4KXL on the dGPU was slower than IQ4XS on the APU alone; after hours of tuning, IQ4XS matched APU token generation with a 2-3x prompt-processing boost and a modest context bump.
Many published dGPU speedups were run at low context, he notes. There is a real tradeoff: offload more layers to the dGPU at the cost of context size/quality, or keep a higher-quality context with fewer dGPU layers. Full Docker command and models.ini settings (tensor-split, cache-type, batch sizes) are included.
More from Infra
- Apple's new Mac mini may arrive before September; enough RAM for local AI worth the cost — Scobleizer · 2026-08-25
- Australia projects data center electricity use will soar nearly 600% by 2036 — Polymarket · 2026-08-25
- Nvidia tells big customers AI chip prices are rising over 15% — emmanuelvivier · 2026-08-25
- AI Supply Chain Faces Bullwhip Effect, HDD Prices Surge — AccBalanced · 2026-08-25
- From Notebook to Production: A 15-Day MLOps Learning Roadmap — _jaydeepkarale · 2026-08-25
- Debunking Data Center Myths: Water, Power, Taxes, and Land Use — AndyMasley · 2026-08-25