Running Qwen3.8-Flash-Next 125B MoE on a 12GB RTX 4070 at ~20 tok/s with MTP

carteakey · reddit · 2026-09-15

The author runs Qwen3.8-Flash-Next (125B-A6B MoE + 51B n-gram table) on RTX 4070 12GB + 64GB DDR5-5600 + Gen4 NVMe under Linux, going from 6 tok/s to nearly 20 tok/s. They claim it beats the 27B dense model on most tasks, ideal for low-VRAM/high-RAM configs.

What worked:

Original post →

More from Infra

Infra channel →