llama.cpp Expert-Pool Fork Boosts Qwen MoE to 17 t/s on a Single MI50 16GB

Atretador · reddit · 2026-09-18

A llama.cpp fork (moe-expert-pool) fixes the admission-budget issue that silently disables expert offloading on 16GB cards, letting a Qwen MoE run on a single MI50. Benchmarks: stock CPU MoE path 11.76 t/s vs 16.39 t/s with 40 expert-cache slots (61.8% hit rate), 17.60 t/s warm, peak 19.8 t/s observed. Full run script with all parameters shared.

Original post →

More from Infra

Infra channel →