ACL 2020 Paper: Reordering Transformer Sublayers Improves Performance

giffmana · x · 2026-07-31

In response to recent discussions about modifying Transformer architectures—such as replacing attention layers with MLPs—a researcher pointed out a relevant paper published seven years ago.

Titled Improving Transformer Models by Reordering their Sublayers (ACL 2020) by Ofir Press et al., the research investigates the ordering of self-attention and feed-forward sublayers. The authors found that breaking the traditional interleaved pattern and adopting a "sandwich" structure—with more self-attention at the bottom and more feed-forward layers at the top—improves perplexity on language modeling benchmarks at no extra parameter, memory, or training cost.

Related event: Researcher Questions Need for Linear Attention, Proposes Minimal MLP Baseline(6 posts)→

Original post →

More from Research

Research channel →