Custom llama.cpp Branch Adds Expert Expansion for MoE Models
Specific-Tax-6700 · reddit · 2026-09-07
A developer built a custom branch of llama.cpp (with help from GLM 5.3 Flash) that supports expert expansion for MoE models. Tested only on Metal so far, and it works better than the author's earlier DS4 version.
The branch is open-sourced with docs, and the author is looking for feedback from other platforms and different models.
- Docs: moe-expansion branch
Related event: Dev builds llama.cpp fork enabling MoE expert expansion(2 posts)→
More from Infra
- AI meme: turning off the tap while brushing teeth to "conserve water for the datacenter buildout" — EigenGender · 2026-09-07
- AMD MI355X beats NVIDIA B300 on tokens-per-dollar in AgentX — AccBalanced · 2026-09-07
- RandKV ships as pip-installable random KV-cache eviction for Transformers, reports honest negative perf results — atease01 · 2026-09-07
- Offline village AI: $5,000 budget to build a local LLM machine for basic Q&A, seeking GPU advice — Potential_Low_1183 · 2026-09-07
- Developer earns just $2.60 per cycle running a bot on OpenAI-subsidized tokens — TheMoonMidas · 2026-09-07
- Interactive Speculative Decoding Tutorial for NeurIPS Explains When It Stays Lossless — Madisonkanna · 2026-09-07