68GB Qwen MoE at 21 tok/s on RTX 3060 + 16GB RAM, bit-exact via new --moe-direct-io

zyxciss · reddit · 2026-10-10

A Redditor implemented --moe-direct-io in llama.cpp, running the 68GB Qwen 3.8 Flash Next (125B MoE, 512 experts, top-10 routing) IQ2XS build at 20-21 tok/s (24+ warm) on an RTX 3060 12GB + 16GB DDR4 machine — a 10-15x speedup over stock llama.cpp's 1.4-2.1 tok/s.

Key details:

A genuine breakthrough for running large MoE models on low-RAM consumer hardware.

Original post →

More from Infra

Infra channel →