Running Qwen3.8-Flash-Next 125B MoE on a 12GB RTX 4070 at ~20 tok/s with MTP
carteakey · reddit · 2026-09-15
The author runs Qwen3.8-Flash-Next (125B-A6B MoE + 51B n-gram table) on RTX 4070 12GB + 64GB DDR5-5600 + Gen4 NVMe under Linux, going from 6 tok/s to nearly 20 tok/s. They claim it beats the 27B dense model on most tasks, ideal for low-VRAM/high-RAM configs.
What worked:
- AtomicChat's 4.27 bpw GGUF quant
- N-gram SSD offloading (lazy-mode)
- --fit on --fit-target 512 for auto param selection
- Master branch: 19.35 t/s; with MTP PR #28243 + Daniel Han's 1.78GB shared-Q4KM compact head and -ncmoe 45, acceptance hits 77-96% and every tested task breaks 20 t/s (peak 20.65)
- Caveat: with low VRAM, MTP gains are limited since layers must be sacrificed for the MTP head
- Prompt processing is still slow at 300-350 tok/s
More from Infra
- Researcher open-sources theseus, a human-language architecture research framework — pratyusha_PS · 2026-09-15
- Banning data centers to save the world? A 500-year history lesson says otherwise — thursdai_pod · 2026-09-15
- Forking Running Machines: Clone a Live Doom Session into 32 Realities in One Second — bigaiguy · 2026-09-15
- Zuckerberg publicly asks Nvidia for a short-lead-time DGX Station for local AI — beffjezos · 2026-09-15
- BIS annual report picked apart: no codified H20 rule, Entity List stalled, loopholes open — ohlennart · 2026-09-15
- Apple's iOS 27 on-device AFM 3 has 20B params, activating only 1-4B — rxwei · 2026-09-15