Qwen 3.6 35b slower on 7900XTX than 3060Ti? User investigates Vulkan config

Jebbyk1 · reddit · 2026-08-25

A user found that the inference speed of the Qwen 3.6 35b (a3b q4-k-m) model on an AMD 7900XTX (20t/s, 100% load) was lower than on an NVIDIA 3060Ti (30t/s, 50% load). Neither Vulkan nor ROCm in a Linux llama.cpp environment resolved the performance regression, despite the theoretical expectation that the 7900XTX should be 2-3 times faster.

Original post →

More from Infra

Infra channel →