Qwen 3.6 35b slower on 7900XTX than 3060Ti? User investigates Vulkan config
Jebbyk1 · reddit · 2026-08-25
A user found that the inference speed of the Qwen 3.6 35b (a3b q4-k-m) model on an AMD 7900XTX (20t/s, 100% load) was lower than on an NVIDIA 3060Ti (30t/s, 50% load). Neither Vulkan nor ROCm in a Linux llama.cpp environment resolved the performance regression, despite the theoretical expectation that the 7900XTX should be 2-3 times faster.
More from Infra
- Token Pooling Can Offset RAG Costs for Long Documents — antoine_chaffin · 2026-08-25
- Running OpenCode inside Durable Objects eliminates the need for a full sandbox — craigsdennis · 2026-08-25
- QuixiAI pushes back on "NVFP4 is for Blackwell" claims — QuixiAI · 2026-08-25
- Raja Koduri's OXMIQ to help build 2GW renewable-powered AI compute platform in India — RajaXg · 2026-08-25
- AM Intelligence confirms binding order for 9,000 Nvidia Rubin NVL72 systems in Hyderabad — RajaXg · 2026-08-25
- xAI's Memphis Data Center Generates Over $100M in Local Taxes, Creates Jobs — robleclerc · 2026-08-25