E-MoE Improves Few-Step Diffusion LMs Using Expert Routing as Shared Latent
Arseny Ivanov · hf · 2026-10-02
Masked diffusion models (MDMs) typically factorize their reverse process over positions, limiting few-step generation quality. Existing continuous Gaussian-latent approaches rely on VAEs and are prone to posterior collapse.
E-MoE instead builds the reverse process as a mixture of factorized distributions over a discrete shared latent defined by the expert-routing decisions of a Mixture-of-Experts backbone, without increasing active parameters over the factorized baseline. It improves few-step generation on synthetic multimodal benchmarks, binarized MNIST, and LM1B.
More from Research
- RAG fixes worked on test questions but not held-out ones: overfitting on 790 clinical PDFs? — Overall_Judge8086 · 2026-10-02
- Critique: Decision Models Leave 10-15% Performance on the Table and Can't Return Evidence — joecole · 2026-10-02
- USTC's PhysVista benchmark exposes wide gap between VLM visual recognition and physical understanding — ustc · 2026-10-02
- NEEDLE: training-free backdoor removal for LLMs drops code-injection attack success to 0% — locailabs · 2026-10-02
- Latent-Foresight: end-to-end latent world models beat two-stage pipelines on future scene prediction — Efstathios Karypidis · 2026-10-02
- Technion: generalization is stability, not accuracy — cross-dataset variance can reverse LLM rankings — Technion · 2026-10-02