Running Qwen 27B on 3x 2080Ti at 55tps: Optimal Config Shared
AccountGotLocked69 · reddit · 2026-08-01
A developer achieved 55 tokens/s running the Qwen 27B Q5 model using llama.cpp on a rig with 3x RTX 2080 Ti (11GB) and a Threadripper 3970X.
They shared the optimal launch parameters, highlighting key optimizations:
- Enabling flash-attn (flash attention)
- Compressing KV cache to q80 precision
- Using tensor split mode and disabling mmap
- Enabling draft-mtp for speculative decoding (max draft of 3)
More from Infra
- DeepSeek V4 Flash local benchmark nearly matches top frontier models from 5 months ago — joorklee · 2026-08-01
- New Method Pre-routes MoE Layers to Optimize I/O for Edge Streaming — dai_app · 2026-08-01
- 5TB of Data Stored on a Tiny Glass Slab Marks Microscopic Storage Breakthrough — TansuYegen · 2026-08-01
- antirez Enables Lossless MXFP4 Local Inference for DeepSeek v4 Flash — antirez · 2026-08-01
- Open-Source Engine 'Waste' Runs Kimi K3 on Just 29GB of RAM — galapag0 · 2026-08-01
- DeepSeek-V3 Trained With Only 180K GPU-Hours, Slashing MoE Compute Costs — teortaxesTex · 2026-08-01