Running Qwen 3.6 35B on a Budget AMD Radeon 7600: Optimized to 21 token/s

Sweaty_Perception655 · reddit · 2026-08-13

A developer on r/LocalLLaMA shared their performance benchmarks and optimization process for running the Qwen 3.6 35B A3B-Q80 GGUF model locally using a budget AMD graphics card (Radeon 7600).

Hardware & Environment: AMD Radeon 7600 GPU, paired with 64GB DDR4 RAM and a Ryzen 5600 processor. Running on Ubuntu via llama.cpp.

Optimization: By rebuilding llama.cpp to support ROCm 6.1.4, generation speed increased from the initial 18 tokens/s to 19+ tokens/s. After maximizing VRAM overclocking via LACTL, the speed reached a stable 21 tokens/s.

Weird Bug: The author discovered a strange rendering bug: if the llama.cpp token generation output window is actively watched, the speed drops to 13 tokens/s; minimizing the window pushes the speed back up to 21 tokens/s.

Original post →

More from Infra

Infra channel →