First PhD paper accepted at NeurIPS 2026: Sparse layers key to scaling looped LMs
burny_tech · x · 2026-09-27
USC PhD student Ryan Lee's first paper, "Sparse Layers are Critical to Scaling Looped Language Models," is accepted at NeurIPS 2026. The paper compares standard transformers, MoE models, and looped architectures, with two main findings:
- Looped-MoE scales better than the standard baseline, while dense looped models do not. The reason is routing divergence across loops: different experts are activated on each pass through the same shared layers, recovering expressivity without extra parameters.
- Loop boundaries are superior early-exit points: since each loop ends with the layers that produce the final output, early exits at loop boundaries yield better compute-quality trade-offs than in standard models.
The authors conclude that a Looped-MoE model with early exits can beat standard transformers at scale while delivering significant memory and inference savings with minimal quality loss — a step toward compute-adaptive LMs.
More from Models
- Opus 5.5 given a creative prompt plus Midjourney overnight产出 draws Elon Musk's "Wow" — tetsuoai · 2026-09-27
- One persona tweak made ChatGPT say 'goblins' 4,000% more — caught on Reddit before OpenAI noticed — victor_explore · 2026-09-27
- MiMo-V2.6 listing hints at 5 models: 1T, 311B and 9B visible, two more unknown — jacek2023 · 2026-09-27
- antirez: Good programmers failing with GPT 6 Astra points to a different skill set — antirez · 2026-09-27
- Opus 5.5 and GPT-6 shipped 101 minutes apart, both cheaper — airesearch12 · 2026-09-27
- Hands-on: Opus 5.5 nails frontend consistency; Astra still wins reasoning — haider1 · 2026-09-27