ByteDance's SplitMoE breaks the uniformity trap to scale video diffusion MoE models
ByteDance · hf · 2026-09-30
ByteDance identifies a 'uniformity trap' in visual MoEs for video diffusion: token-wise routing with uniform expert-usage regularization scatters coherent patches across experts, causing fragmentation and structural distortion. SplitMoE bifurcates the expert pool into semantic experts (high-level abstraction) and generic experts (residual visual info), using prototype-guided routing and pull-push regularization. Under equal activated-parameter budgets, it outperforms load-balanced MoEs in convergence, routing coherence, and video quality, and reveals an emergent coarse-to-fine denoising logic.
More from Multimodal
- Four imaginary tokens for Midjourney v8.2 produce memory ghosts and bone echoes — LudovicCreator · 2026-09-30
- Hyper-personalized music is BS: music is culture and inherently social, argues developer — jordiponsdotme · 2026-09-30
- NUS Proposes StoryEngine: A State-Grounded Agentic Framework for Coherent Long-Form Video Storytelling — NationalUniversityofSingapore · 2026-09-30
- One Year of Local Image Generation: Why Civitai and ComfyUI Both Fall Short — BenDLH · 2026-09-30
- Opus made a launch video for Violetto 1B in 50 minutes amid zero media coverage — tensorqt · 2026-09-30
- LoRA adapters break on distilled video models, long post explains why — burkov · 2026-09-30