Running 180B Qwen3.8-Flash-Next at 40-50 t/s on 64GB RAM + 16GB VRAM with Strata

danamir_ · reddit · 2026-10-03

A detailed hands-on report of running Qwen3.8-Flash-Next (180B total params: 125B with 6B activated, plus 51B n-gram embedding and 4B MTP) locally via Strata. On a Windows 10 machine with 64GB RAM and a 5070Ti (16GB VRAM), the IQ3XXS quant runs at 40-50 t/s with up to 256K context.

Key points:

Original post →

More from Infra

Infra channel →