Running 180B Qwen3.8-Flash-Next at 40-50 t/s on 64GB RAM + 16GB VRAM with Strata
danamir_ · reddit · 2026-10-03
A detailed hands-on report of running Qwen3.8-Flash-Next (180B total params: 125B with 6B activated, plus 51B n-gram embedding and 4B MTP) locally via Strata. On a Windows 10 machine with 64GB RAM and a 5070Ti (16GB VRAM), the IQ3XXS quant runs at 40-50 t/s with up to 256K context.
Key points:
- Strata handles model loading/unloading automatically between prompting sessions (2 min load), letting users run ComfyUI in between for image work.
- Adding the mmproj enables vision capabilities; 32-64K context suffices for use as an image assistant.
- Alternative: run directly on llama.cpp with --cpu-moe, but expect roughly half the speed.
More from Infra
- GKE adds CPU Startup Boost: faster pod starts without over-provisioning — rseroter · 2026-10-03
- Huawei claims Ascend overtook Nvidia in China share; supply, not demand, is the bottleneck — teortaxesTex · 2026-10-03
- Turning an iPhone into a second GPU for a MacBook: 44% faster prefill on Qwen 27B — StayLameBro · 2026-10-03
- Micron CEO: memory supply will be much tighter in 2027-2028 than 2026 — dankvr · 2026-10-03
- mamf-finder adds FP8/MXFP4/NVFP4 support for real GPU TFLOPS benchmarking — StasBekman · 2026-10-03
- Measured on B200: nvfp4 is ~9% more efficient than mxfp4 with higher accuracy — pick nvfp4 on Blackwell — StasBekman · 2026-10-03