llama.cpp PR fuses shared experts into MMVQ, speeding up MoE inference for select architectures

jacek2023 · reddit · 2026-10-03

Developer am17an opened Pull Request #29184 on ggml-org/llama.cpp that fuses shared experts into the MMVQ quantization path to speed up MoE inference. The author notes the speedup only applies to some MoE architectures, such as Qwen 35B A3B, not all MoE models. A notable performance improvement for users running quantized MoE models locally.

Original post →

More from Infra

Infra channel →