MoE-to-Dense scores NeurIPS 2026 oral: distilling dense models from MoEs without retraining
Kangwook_Lee · x · 2026-09-30
The paper "MoE-to-Dense" has been selected for an oral presentation at NeurIPS 2026 in Sydney.
Motivation: almost all flagship models are now MoEs, but smaller models still prefer dense architectures since they target memory-constrained scenarios where total parameter count matters. The paper asks whether a pre-trained MoE can be leveraged to produce dense models without training from scratch — offering an efficient path to compact dense models distilled from large MoEs.
More from Research
- Diffusion Models Tutorial Accepted to NeurIPS 2026 Alongside 7 Paper Acceptances — mittu1204 · 2026-09-30
- AMB3R-SLAM: Kilometer-Scale Real-Time SLAM on One Consumer GPU, Cutting ATE by 70% — rsasaki0109 · 2026-09-30
- Prefix-Reuse FLOPs: new metric exposes hidden cost of arbitrary context edits in LLM serving — RulinShao · 2026-09-30
- Tencent Hunyuan releases ExplorationBench to measure how AI systems explore — TencentHunyuan · 2026-09-30
- Fully open MolmoAct 2 tops independent robotics benchmark LIBERO-MAX on dynamic robustness — DJiafei · 2026-09-30
- CompVis improves Distributional Diffusion Models: 4.48 FID at 4 steps on ImageNet — CompVis · 2026-09-30