E-MoE Improves Few-Step Diffusion LMs Using Expert Routing as Shared Latent

Arseny Ivanov · hf · 2026-10-02

Masked diffusion models (MDMs) typically factorize their reverse process over positions, limiting few-step generation quality. Existing continuous Gaussian-latent approaches rely on VAEs and are prone to posterior collapse.

E-MoE instead builds the reverse process as a mixture of factorized distributions over a discrete shared latent defined by the expert-routing decisions of a Mixture-of-Experts backbone, without increasing active parameters over the factorized baseline. It improves few-step generation on synthetic multimodal benchmarks, binarized MNIST, and LM1B.

Original post →

More from Research

Research channel →