ACL 2020 Paper: Reordering Transformer Sublayers Improves Performance
giffmana · x · 2026-07-31
In response to recent discussions about modifying Transformer architectures—such as replacing attention layers with MLPs—a researcher pointed out a relevant paper published seven years ago.
Titled Improving Transformer Models by Reordering their Sublayers (ACL 2020) by Ofir Press et al., the research investigates the ordering of self-attention and feed-forward sublayers. The authors found that breaking the traditional interleaved pattern and adopting a "sandwich" structure—with more self-attention at the bottom and more feed-forward layers at the top—improves perplexity on language modeling benchmarks at no extra parameter, memory, or training cost.
More from Research
- Critique of entropy decomposition in AI safety research — AdaptiveAgents · 2026-08-25
- Scaffold CoT: A 4M-Example Dataset Teaching Small Models to Think in a Fixed Structure — Saraozte01 · 2026-08-25
- TiDE-Ab paper introduces time-dependent guidance for antibody design — DaveJuergens · 2026-08-25
- Free open-source interactive explainer on World Models — Dooraven · 2026-08-25
- DiffSynth open-sources MiniMax-H3 LoRA training adapter and dataset — bdsqlsz · 2026-08-25
- Alibaba Releases Swift-Image: A Compact 6B Unified Text-to-Image and Editing Model — HaktanSuren · 2026-08-25