Optimizing Qwen 3.8 Flash Next: Improving speeds on 64GB VRAM setup

Jorlen · reddit · 2026-09-01

A user discusses optimizing Qwen 3.8 Flash Next (Q4 quantized) on a dual R9700 (64GB VRAM) + 32GB RAM setup. Current pre-fill is 180 t/s with generation at 18 t/s. Limited by 32GB system RAM preventing full model loading into memory, the user relies on MMAP. With flash-attn enabled and KV cache adjusted, the user seeks further methods to increase throughput.

Original post →

More from Infra

Infra channel →