MoE-to-Dense scores NeurIPS 2026 oral: distilling dense models from MoEs without retraining

Kangwook_Lee · x · 2026-09-30

The paper "MoE-to-Dense" has been selected for an oral presentation at NeurIPS 2026 in Sydney.

Motivation: almost all flagship models are now MoEs, but smaller models still prefer dense architectures since they target memory-constrained scenarios where total parameter count matters. The paper asks whether a pre-trained MoE can be leveraged to produce dense models without training from scratch — offering an efficient path to compact dense models distilled from large MoEs.

Original post →

More from Research

Research channel →