Dual RX 6800 running Qwen3 27B: Linux + ROCm beats Windows Vulkan, 45 tok/s TG
SaGa31500 · reddit · 2026-10-06
The author ran Qwen3 27B q6k on dual RX 6800/6800XT (gfx1030) GPUs, hitting a Windows 11/Vulkan bug that halved prompt processing before switching to Linux + ROCm.
- Windows/Vulkan: 35 tok/s TG, 180 tok/s PP with MTP on (bug-limited); 20 tok/s TG, 360 tok/s PP without MTP; 15 tok/s TG with tensor split.
- Linux/ROCm: 45 tok/s TG and 450 tok/s PP with tensor split + MTP; small ub (256) improves PP.
- The author asks which inference engines and parameters to try next, and for comparable dual-gfx1030 numbers.
More from Infra
- Dev's custom vLLM patch runs DiffusionGemma on DGX Spark, edging out the commercial API — vllm_project · 2026-10-06
- Cloudflare launches cf, an agent-first CLI to query observability data via the API — dinasaur_404 · 2026-10-06
- Building a local LLM agent stack on a 128GB Mac Studio: Reddit thread weighs inference layer options — DrainBramage · 2026-10-06
- Polars 2.0 ships: streaming by default, SQL first-class, beats DuckDB on TPC-H — banteg · 2026-10-06
- Qwen3-Next-80B on a 3090 with 16GB RAM: ~3x decode speedup via MoE expert substitution — Zestyclose_Reality15 · 2026-10-06
- Disco a dangerous short after HBM despec news; H2 China demand may flip — zephyr_z9 · 2026-10-06