20+ Hand-Tuned ROCm Kernels Nearly Double Qwen 27B Throughput on 4x 7900XTX

NoFee9147 · reddit · 2026-10-07

A user, working with Claude, implemented 20+ fixes in ROCm and llama.cpp server to run Qwen3.8-27B Q8 on 4x Radeon 7900XTX (Lenovo P620, PCIe 4.0 x16), lifting code decode from 54-61 to 96-110 tok/s and 51K prefill from 1454 to 1816 tok/s. Highlights:

Configs and code to be pushed to GitHub if there's interest.

Original post →

More from Infra

Infra channel →