Running Qwen 3.5 35B at 18 token/s on RTX 5080 Setup
Sweaty_Perception655 · reddit · 2026-08-10
A developer successfully ran the Qwen 3.5 35B A3B-Q80 gguf model at 18 token/s using llama.cpp on Ubuntu. The hardware setup includes a Radeon 7600 GPU paired with 64GB DDR4 RAM and a Ryzen 5600 CPU.
The specific configuration includes offloading 37 MoE layers to the CPU (--n-cpu-moe 37), disabling mmap, using Q8 quantization for context (-ctk q80 -ctv q80), enabling Flash Attention (-fa 1), and setting a context length of 9000.
More from Infra
- Report: Nvidia Qualifies 300mW Lasers, Buys Bulk of Supply — zephyr_z9 · 2026-08-10
- Dual-GPU Optimization Speeds Up MiniMax-H3 Video Generation 8x — multimodalart · 2026-08-10
- PyTorch DevLog: Why You Should Never Free Pinned Memory — ezyang · 2026-08-10
- Rust Linear Algebra to Wasm Achieves 6x Browser AI Performance Boost — doodlestein · 2026-08-10
- France's 10GW Power Surplus Could Yield €350B Annually via AI Datacenters — emmanuelvivier · 2026-08-10
- Won 5th Place in GPU Mode with Coding Agents, No CUDA Background — tokenbender · 2026-08-10