Sparse Universal Transformer: SMoE makes parameter-sharing transformers scalable (EMNLP 2023)
Yikang_Shen · x · 2026-09-04
In a discussion about looped transformers and adaptive per-token computation, Yikang Shen points out this is the Universal Transformer idea and shares his earlier paper, Sparse Universal Transformer (EMNLP 2023, with Shawn Tan, Aaron Courville, Chuang Gan):
- UT shares parameters across layers, is Turing-complete under assumptions, generalizes compositionally better than vanilla transformers, but scales poorly due to compute/memory costs
- SUT uses Sparse Mixture of Experts (SMoE) to cut UT's computational complexity while keeping parameter efficiency and generalization
- Experiments show strong results on formal-language tasks (logical inference, CFQ) plus impressive parameter/compute efficiency on benchmarks like WMT'14
More from Research
- Mol-JEPA: A Multimodal JEPA Foundation Model for Molecules, One Year in the Making — TerribleAntelope9348 · 2026-09-04
- Two Years On: Six Guidelines for Making Research Impact via Open-Source in AI — lateinteraction · 2026-09-04
- GPT-6 Astra claims SOTA on ARC-AGI-3 at 66%, up from Sol's 8% — teortaxesTex · 2026-09-04
- Steering Qwen along a grader-vs-human dimension oddly shifts its personality — voooooogel · 2026-09-04
- CMU's AI Reviewer Beats Best Human Reviewer, Featured by Science — AkariAsai · 2026-09-04
- Chollet: ARC-AGI-4 lands Q1 2027, and solving ARC-3 is not AGI — fchollet · 2026-09-04