Mixture of Heterogeneous Grouped Experts for Language Modeling
Zhicheng Ma, Xiang Liu, Zhaoxiang Liu, Ning Wang, Yi Shen, Kai Wang, Shuming Shi, Shiguo Lian
cs.CL, cs.AI, cs.LG
2026-04-25
MoHGE groups experts by width with two-stage routing. It matches MoE scores with ~16% fewer total params, ~25% fewer activated expert params, and even GPU load.
In a standard MoE, every expert has the same width. Easy tokens and hard tokens pay the same compute bill. That is wasted capacity in production inference: a large share of tokens never need the widest expert.
Heterogeneous experts try to match width to difficulty. MoDSE pushes routing probabilities toward uniformity, so tokens often miss the expert that actually fits and small experts stay idle. HMoE sketched a hybrid homogeneous-heterogeneous layout, but mixing widths across GPUs leaves some devices saturated and others idle. Training throughput and scale-out both stall.
Two constraints have to hold together: capacity should track token difficulty, and GPU work still has to stay even.
MoHGE buckets experts into groups. Width is shared inside a group and stepped across groups. Group i uses hidden size Wi = 2 Wbase - W{Ng-i}. Small groups cover simple patterns; larger groups take harder context.
Routing runs in two stages. A group gate scores each token with a sigmoid against group centroids and keeps the top-Kg groups. An expert gate then softmaxes inside those groups, multiplies by the group score, takes a global top-Ke, and renormalizes. Locking groups first, then experts, yields more combinations than a fully heterogeneous set where each expert has a unique size, and compute stays inside the chosen groups.
Large experts win routing if left unconstrained. Group-Wise Auxiliary Loss adds a tiny penalty αG scaled by each group's parameter count Wi, so the gate trades cross-entropy against parameter cost and prefers a smaller group when it is enough.
GPU balance is placement plus a second loss. All-size Group-decoupling Allocation packs the i-th expert from every group onto the same GPU, so every device holds the same total expert parameters. Intra-Group Experts Auxiliary Loss takes DeepSeekV2's global expert-balance term and applies it inside each group, equalizing how often experts in an active group are picked. Once groups are internally even, the All-size packing makes devices even too.
Main numbers follow OpenCompass, zero-shot or few-shot, averaged over three runs.
| Method | Total params | Activated expert params | MMLU | GSM8K | MATH | TriviaQA |
| Dense | 0.807B | – | 26.36 | 2.79 | 1.33 | 34.98 |
| MoE-3B | 3.3614B | 0.376B | 26.22 | 3.03 | 1.34 | 39.16 |
| MoHGE-3B | 2.821B | 0.295B | 26.41 | 4.02 | 1.36 | 39.20 |
| Dense | 1.672B | – | 30.78 | 4.62 | 6.82 | 50.26 |
| MoE-14B | 16.760B | 1.191B | 31.18 | 4.92 | 7.30 | 51.77 |
| MoHGE-14B | 14.122B | 0.843B | 31.62 | 5.76 | 9.42 | 52.69 |
The abstract claims about 20% fewer total parameters. Table 1 is closer to 16%: 3.3614B to 2.821B at 3B, 16.760B to 14.122B at 14B. Activated expert parameters drop by about a quarter, 0.376B to 0.295B and 1.191B to 0.843B. MoHGE matches or slightly beats the matched MoE on all seven tasks. At 14B, MATH moves from 7.30 to 9.42, GSM8K from 4.92 to 5.76, MMLU from 31.18 to 31.62. At 3B, GSM8K rises from 3.03 to 4.02, a larger relative bump with scores still in single digits. PIQA is tied at 49.08.
On downstream eval wall time, MoHGE is slightly faster than the matched MoE on most tasks (14B MMLU 18.86h vs 19.27h). GSM8K prefers larger groups, so 3B takes 0.86h vs 0.84h for MoE. The dense models are faster still: they skip routing, and their total size is about the MoE activated size.
For the 14B GPU study, the i-th expert of each group sits on GPU i. Token share per group per GPU sits near 12.5%, with intra-group standard deviation about 0.002 to 0.003. A 3B reproduction of MoDSE and HMoE has MoHGE slightly ahead on MMLU (26.41 vs 26.27 / 26.34) and GSM8K (4.02 vs 3.71 / 3.94). HMoE is close on quality and unbalanced on GPUs.
Routing follows the intended difficulty split. By training-corpus frequency, Top 1K tokens go to the smallest Group 1 16.3% of the time and to the largest Group 8 only 9.7%; Beyond 10K tokens fall to 11.3% on Group 1 and rise to 13.3% on Group 8. The same tilt appears when tokens are binned by perplexity. Drop the Group-Wise Auxiliary Loss and mass shifts back onto large groups.
This is an industrial, incremental systems result: make heterogeneous MoE deployable with even GPU loads. Teams chasing inference cost can reuse the two-level router and the All-size pack to send easy tokens to small groups, without a unique width per expert.
It does not contest frontier-scale production MoE. At 3B and 14B total size on this OpenCompass slice, the setup yields about one-sixth fewer total parameters, about one-quarter fewer activated expert parameters, no score drop, and flat GPU routing.
The main text omits pretraining token count, training steps, group count Ng, Kg and Ke, and the actual αG / αE values. Ablations sit in the appendix. How the 3B and 14B MoE baselines were matched on depth, width, and activated experts is not fully specified.
Absolute scores are low. MoHGE-14B reaches 5.76 on GSM8K and 9.42 on MATH. That reads as a short-train or small-corpus relative comparison, not a production checkpoint. The abstract's "about 20%" cut does not match the 16% in Table 1, which will get rounded up in retellings.
The HMoE baseline is a self-reproduced Top-P variant; fidelity to the original paper is unverifiable from the main text. All-size packing assumes intra-group balance. Inference traffic that saturates one group could unbalance devices again. The paper reports eval-time routing, not a load test under skewed traffic.