First PhD paper accepted at NeurIPS 2026: Sparse layers key to scaling looped LMs

burny_tech · x · 2026-09-27

USC PhD student Ryan Lee's first paper, "Sparse Layers are Critical to Scaling Looped Language Models," is accepted at NeurIPS 2026. The paper compares standard transformers, MoE models, and looped architectures, with two main findings:

The authors conclude that a Looped-MoE model with early exits can beat standard transformers at scale while delivering significant memory and inference savings with minimal quality loss — a step toward compute-adaptive LMs.

Original post →

More from Models

Models channel →