China Unicom's Heterogeneous Grouped MoE Cuts Parameters 20%, Wins ACL 2026 Spot
量子位 · wechat · 2026-10-04
China Unicom's AI research team introduced MoHGE, a Mixture of Heterogeneous Grouped Experts architecture accepted to ACL 2026 with open-sourced code. It addresses homogeneous MoE waste and GPU load imbalance via grouped heterogeneous experts with two-level routing, parameter-based auxiliary loss steering simple tokens to smaller experts, and an all-size group-decoupling allocation strategy. At 3B and 14B scales it cuts total parameters by 20% and activated parameters by 25% versus equal-accuracy MoE baselines, while matching or beating accuracy on MMLU, MATH, and GSM8K with near-perfect GPU load balance.
More from Infra
- UBS Sees Global Rack Capacity Hitting 104.5 GW by 2030, Half Going to Nvidia — AccBalanced · 2026-10-04
- Musk: xAI to hit 10GW of compute by end of next year, sees AI inference moving to space — beffjezos · 2026-10-04
- BF16 rounding breaks a conservation law, blowing up FlashAttention gradients late in training — HongyiWang10 · 2026-10-04
- Getting PyTorch CUDA training running on BC-250 boards, captured as an image — redfoxkiller · 2026-10-04
- Huawei 950 super-node claims seamless scaling from 550B/1.6T up to 10T-class models — teortaxesTex · 2026-10-04
- The curse of 64GB RAM: Strata pushes local Qwen3.8-Flash-Next to 60 t/s but hogs system memory — Cautious_Chicken_604 · 2026-10-04