Running Qwen3.8-Flash-Next on 2x3090: experts to RAM, 51B n-gram table on NVMe

jbro1985 · reddit · 2026-08-29

A Reddit user details deploying Qwen3.8-Flash-Next (125B MoE + 51B n-gram table) on 2x RTX 3090 + 96GB DDR5, with experts offloaded to RAM and the n-gram table mmap'd to NVMe.

Llama.cpp

vLLM

Original post →

More from Infra

Infra channel →