Local Qwen3 8B Flash on RTX 6000 Pro: 2000 tps prefill but only 40 tps decode
AppealSame4367 · reddit · 2026-09-04
A user runs Qwen3 8B Flash on a rented RTX 6000 Pro via ikllama (3 slots, 200k context, IQ4 quant, q8 KV cache): 2000 tps prefill but only 40 tps decode for a single request, dropping to 10-20 tps decode with 2-4 parallel requests. vLLM recipes didn't help; they're asking for better configs for 2-4 slots at q4+ quants.
More from Infra
- Morgan Stanley: 20% of Star Market IPOs now target China's 'chokepoint' tech — pstAsiatech · 2026-09-04
- DeepSeek plans 160,000-chip Huawei Ascend 950DT cluster in Inner Mongolia, Bloomberg reports — kimmonismus · 2026-09-04
- Zeiss executive: China about 15 years behind on EUV lithography tools — pstAsiatech · 2026-09-04
- Minisforum MS-S1 MAX P495 listing hints at €7999 price, double the original MS-S1 MAX — fairydreaming · 2026-09-04
- First community MLX 4-bit benchmarks for K2-Horizon-MoVA-36B hit 49.1 tok/s locally — DerTomsn · 2026-09-04
- TSMC doubles equipment forecast in six months, builds 20 fabs to chase AI demand — firstadopter · 2026-09-04