llama.cpp Expert-Pool Fork Boosts Qwen MoE to 17 t/s on a Single MI50 16GB
Atretador · reddit · 2026-09-18
A llama.cpp fork (moe-expert-pool) fixes the admission-budget issue that silently disables expert offloading on 16GB cards, letting a Qwen MoE run on a single MI50. Benchmarks: stock CPU MoE path 11.76 t/s vs 16.39 t/s with 40 expert-cache slots (61.8% hit rate), 17.60 t/s warm, peak 19.8 t/s observed. Full run script with all parameters shared.
More from Infra
- NanoGPT Speedrun hits new 68.0s record with 96-dim QK and packed FP8 attention — kellerjordan0 · 2026-09-18
- Hedge fund CIO: Anthropic burns maybe 80% less per token than OpenAI — rohanpaul_ai · 2026-09-18
- Engineer: companies hide servers from management to dodge forced cloud migrations — irth7 · 2026-09-18
- SK Hynix subsidiary Solidigm plans to build a NAND flash fab in the US — zephyr_z9 · 2026-09-18
- Huawei says China's AI training will shift to Ascend 950DT SuperPoDs starting 2027 — pstAsiatech · 2026-09-18
- Jev playground shows inference and roundtrip latency; EU users pay 120ms extra — DanielLockyer · 2026-09-18