ByteDance's SplitMoE breaks the uniformity trap to scale video diffusion MoE models

ByteDance · hf · 2026-09-30

ByteDance identifies a 'uniformity trap' in visual MoEs for video diffusion: token-wise routing with uniform expert-usage regularization scatters coherent patches across experts, causing fragmentation and structural distortion. SplitMoE bifurcates the expert pool into semantic experts (high-level abstraction) and generic experts (residual visual info), using prototype-guided routing and pull-push regularization. Under equal activated-parameter budgets, it outperforms load-balanced MoEs in convergence, routing coherence, and video quality, and reveals an emergent coarse-to-fine denoising logic.

Original post →

More from Multimodal

Multimodal channel →