Optimizing Qwen 3.8 Flash Next: Improving speeds on 64GB VRAM setup
Jorlen · reddit · 2026-09-01
A user discusses optimizing Qwen 3.8 Flash Next (Q4 quantized) on a dual R9700 (64GB VRAM) + 32GB RAM setup. Current pre-fill is 180 t/s with generation at 18 t/s. Limited by 32GB system RAM preventing full model loading into memory, the user relies on MMAP. With flash-attn enabled and KV cache adjusted, the user seeks further methods to increase throughput.
More from Infra
- Over 4 Petabytes of Models and Datasets Uploaded to Hugging Face in One Week — victormustar · 2026-09-01
- Analyst Weighs Long-Term Impact of Nvidia-MediaTek Deal on Broadcom — BenBajarin · 2026-09-01
- Europe invests €387.8M in LUMI-AI supercomputer with 10x AI capacity — wkmyrhang · 2026-09-01
- SK hynix reportedly considers Intel Foundry for HBM4E base dies — AccBalanced · 2026-09-01
- Guardian proposes off-grid, self-powered datacenters to cut emissions — nordicinst · 2026-09-01
- Compute Wants to Leave Earth: A Manifesto for Orbital Infrastructure — McDonaghMatthew · 2026-09-01