UniMoMo: Shrinking MoE Recommendation Models via Expert Merging

_reachsumit · x · 2026-08-11

UniMoMo introduces an acceleration method for Mixture-of-Experts (MoE) architectures in large recommendation models.

The approach merges experts based on their functional similarity and traffic patterns. This significantly reduces the total number of experts while preserving ranking quality and inference speed, which is crucial for optimizing the deployment and operational costs of massive recommendation systems.

Original post →

More from Research

Research channel →