Looped transformer debate revisited: Sparse Universal Transformer (EMNLP 2023) got there first
Yikang_Shen · x · 2026-09-04
Researcher Yikang Shen points out that today's looped-transformer and per-token adaptive-computation discussions echo the Sparse Universal Transformer (EMNLP 2023). SUT builds on Universal Transformer's parameter sharing, uses Sparse MoE to cut compute/memory costs, and shows stronger compositional generalization on formal-language tasks (logical inference, CFQ) while staying efficient on WMT'14.
More from Research
- Meta's long-context MRCR scores flagged as overfit: 1k samples can lift 60% to 90%+ — eliebakouch · 2026-09-04
- Levin Lab launches platform to train non-neural human cells — drmichaellevin · 2026-09-04
- Model routing cuts LLM errors 46% at same cost, Martian study finds — SucceededMind · 2026-09-04
- matklad on object pools and memory safety: how pooling changes use-after-free effects — jedisct1 · 2026-09-04
- Agent's Last Exam tops out at 59.3% — the benchmark that matters for AI replacing humans — DevToD4 · 2026-09-04
- Has Anyone Tried Feeding All the Bio x ML Datasets to a Single Model? — iskander · 2026-09-04