Custom llama.cpp build pushes 7900XTX to 1600tk/s prefill on Qwen 27B Q8

nasone32 · reddit · 2026-09-08

A Reddit user released a custom llama.cpp build deeply optimized for AMD 7900XTX (single or dual card), targeting usable speeds for Qwen models on consumer ROCm hardware.

Headline results:

Key techniques include a custom HIP allreduce path enabling tensor parallel for chipset-behind cards, optional Q80 compression of inter-card PCIe traffic, --adaptive-mtp, DFLASH2 on tensor parallel, and unmerged upstream PRs (lazy PLE load path +58.88% Flash, GPU MoE expert cache +19.95% decode, etc.).

Tested on Ubuntu 24 with ROCm 7.14. The author offers no support, hoping the patches get merged upstream so the "frankenstein" build can die peacefully.

Original post →

More from Infra

Infra channel →