Running Qwen 3.6 35B on a Budget AMD Radeon 7600: Optimized to 21 token/s
Sweaty_Perception655 · reddit · 2026-08-13
A developer on r/LocalLLaMA shared their performance benchmarks and optimization process for running the Qwen 3.6 35B A3B-Q80 GGUF model locally using a budget AMD graphics card (Radeon 7600).
Hardware & Environment: AMD Radeon 7600 GPU, paired with 64GB DDR4 RAM and a Ryzen 5600 processor. Running on Ubuntu via llama.cpp.
Optimization: By rebuilding llama.cpp to support ROCm 6.1.4, generation speed increased from the initial 18 tokens/s to 19+ tokens/s. After maximizing VRAM overclocking via LACTL, the speed reached a stable 21 tokens/s.
Weird Bug: The author discovered a strange rendering bug: if the llama.cpp token generation output window is actively watched, the speed drops to 13 tokens/s; minimizing the window pushes the speed back up to 21 tokens/s.
More from Infra
- Microsoft's Homegrown AI Chip Effort Shows Signs of Life After Slow Start — pstAsiatech · 2026-08-13
- Why Nvidia Is Trying To Develop the World's Best Open-Source AI Models — pstAsiatech · 2026-08-13
- Saving $40K in Monthly API Costs via Context Caching — NathanWilbanks_ · 2026-08-13
- LTX-Video Local Test: 30-Second Generation Takes 7 Minutes on RTX 5060 — Character_Title_876 · 2026-08-13
- CoreWeave's Profit Surprise Boosts AI Compute Confidence; Quantum Market Heads for Duopoly — TiernanRayTech · 2026-08-13
- ComfyUI Dual-GPU Pitfall: Second GPU Triggers OOM Errors — tricck3zz · 2026-08-13