Running Qwen 3.5 35B at 18 token/s on RTX 5080 Setup

Sweaty_Perception655 · reddit · 2026-08-10

A developer successfully ran the Qwen 3.5 35B A3B-Q80 gguf model at 18 token/s using llama.cpp on Ubuntu. The hardware setup includes a Radeon 7600 GPU paired with 64GB DDR4 RAM and a Ryzen 5600 CPU.

The specific configuration includes offloading 37 MoE layers to the CPU (--n-cpu-moe 37), disabling mmap, using Q8 quantization for context (-ctk q80 -ctv q80), enabling Flash Attention (-fa 1), and setting a context length of 9000.

Original post →

More from Infra

Infra channel →