20+ Hand-Tuned ROCm Kernels Nearly Double Qwen 27B Throughput on 4x 7900XTX
NoFee9147 · reddit · 2026-10-07
A user, working with Claude, implemented 20+ fixes in ROCm and llama.cpp server to run Qwen3.8-27B Q8 on 4x Radeon 7900XTX (Lenovo P620, PCIe 4.0 x16), lifting code decode from 54-61 to 96-110 tok/s and 51K prefill from 1454 to 1816 tok/s. Highlights:
- Comms/dispatch: async mirrored input uploads (24.6→36.3 t/s), per-device dispatch threads, one-shot PCIe P2P AllReplace of RCCL (57µs→9µs).
- MTP decode: fused copies, draft head prefetching only last 2048 prompt tokens (decode 46.6→64.9 t/s), draft buffer cut 457→247 MiB fixing 96K OOM.
- Kernels: mmvq small-K packing, 8-token MoE vector kernel, fused hyper-connection chains (380 fewer kernels/token), thin-F32 prefill kernel (401→21µs).
- MoE/attention: compact expert tile lists (928→405µs), multi-warp routing (99→25µs), QSA sparse attention attending only 2K tokens (67.9→80.6 t/s at 48K).
- Misc: 26.8GB embedding table resident in RAM, MTP graph re-reserve fix, no 64-bit div/mod, split-K router GEMM.
Configs and code to be pushed to GitHub if there's interest.
More from Infra
- 100B decisions in 5 minutes: a layered agent pipeline where Opus only probes 96 hotspots — Arindam_1729 · 2026-10-07
- OpenSBI and Linux now boot on Maxion cores of ET-SOC1 cards, porting done agentically — glenbeer · 2026-10-07
- Samsung's Q3 profit reportedly set to soar 770% YoY as AI demand overwhelms memory supply — Polymarket · 2026-10-07
- openTPU: A One-Person Open-Source AI Accelerator Designed by AI Agents, Running 10 LLMs on an FPGA — ai · 2026-10-07
- Free Online Conference All Day AI Set for Oct 22 With Talks on SLMs and Local Inference — FikoFox · 2026-10-07
- Best $4,000 Local LLM Rig? Weighing R9700s, Strix Halo, and 6x Arc B60 — BinaryGrind · 2026-10-07