China Unicom's Heterogeneous Grouped MoE Cuts Parameters 20%, Wins ACL 2026 Spot

量子位 · wechat · 2026-10-04

China Unicom's AI research team introduced MoHGE, a Mixture of Heterogeneous Grouped Experts architecture accepted to ACL 2026 with open-sourced code. It addresses homogeneous MoE waste and GPU load imbalance via grouped heterogeneous experts with two-level routing, parameter-based auxiliary loss steering simple tokens to smaller experts, and an all-size group-decoupling allocation strategy. At 3B and 14B scales it cuts total parameters by 20% and activated parameters by 25% versus equal-accuracy MoE baselines, while matching or beating accuracy on MMLU, MATH, and GSM8K with near-perfect GPU load balance.

Original post →

More from Infra

Infra channel →