ROCm vs Vulkan on R9700 + Strix Halo: ROCm still wins for DeepSeek, Vulkan closes in
Hrethric · reddit · 2026-09-22
Reddit user ainsoph00 benchmarked ROCm 10.0 vs Vulkan (Mesa 26.2.3 RADV) with llama.cpp 0.4.1-dev on an R9700 + Strix Halo, running DSv4 Flash 0731 and Qwen 3.8 Flash Next after major environment upgrades (kernel, firmware, ROCm 7.14→10, Mesa, and many llama.cpp improvements).
Results:
- DSv4 Flash 0731 (UD-IQ4NL + DSpark drafter): ROCm remains best, but Vulkan has closed the gap — ROCm hit 154.7 tok/s prompt processing and 24.1 tok/s generation on a needle test
- Qwen 3.8 Flash Next (UD-Q4K-XL, no speculative decoding): murkier — ROCm leads prompt processing (566 tok/s), Vulkan leads token generation
Tests used a 34k-word four-needle prompt and the classic Coleridge v Tennyson prompt, with outputs judged by Claude Opus 5: Qwen (Model B) offered a real argument, specific evidence, and genuine insight but several confident factual errors; DeepSeek (Model A) was safer but generic, with interpretive rather than factual errors — B clearly wins for knowledgeable readers, A misleads less for credulous ones.
More from Infra
- Hands-on LLM inference: boosting tokens-per-second with a 31B Gemma model — abhijithneil · 2026-09-22
- rakyll: fully managed platforms fail on composability — bespoke and open source stacks win — rakyll · 2026-09-22
- Local Qwen 27B agent logs into Amazon and buys paper autonomously in one run — fuzhongkai · 2026-09-22
- Qwen Code TUI loses a transcript line on every rows-only shrink, blamed on ink 7.0.3 — SnowCore8 · 2026-09-22
- 60 Minutes: US golf courses use more than twice as much water as data centers — SumitGup · 2026-09-22
- DeltaTensors Stores Fine-Tunes as Weight Deltas: 953MB → 294MB With Minimal Quality Loss — cupheadgamer · 2026-09-22