llama.cpp PR fuses shared experts into MMVQ, speeding up MoE inference for select architectures
jacek2023 · reddit · 2026-10-03
Developer am17an opened Pull Request #29184 on ggml-org/llama.cpp that fuses shared experts into the MMVQ quantization path to speed up MoE inference. The author notes the speedup only applies to some MoE architectures, such as Qwen 35B A3B, not all MoE models. A notable performance improvement for users running quantized MoE models locally.
More from Infra
- AMD to present FlyDSL, an MLIR-native GPU kernel backend for TorchInductor, at PyTorchCon — PyTorch · 2026-10-03
- AMD updates Ryzen AI Developer Platform OS with ROCm 10.0 and Linux 7.2 — Fcking_Chuck · 2026-10-03
- Suhail hails the start of the Vera Rubin era as NVIDIA's next-gen GPU boots up — rickasaurus · 2026-10-03
- GPU-backed loans put GPU earning power and collateral value under scrutiny — AnneliesGamble · 2026-10-03
- Google accused of illegally bulldozing 300 million sq. meters of Finnish forest for AI data centers — Polymarket · 2026-10-03
- SkyRL v0.4 trains 1T-param Kimi K2.7 with RL on just 16 B300 GPUs — casper_hansen_ · 2026-10-03