Dev builds llama.cpp fork enabling MoE expert expansion

A developer created a llama.cpp fork implementing training-free MoE expert expansion, activating more routing experts than the native top-K at runtime, tested with GLM models.

2026-09-07 ~ 2026-09-07 · 2 related posts