Transformer Best Practices Radically Changed Since 2017

ChrisGPotts · x · 2026-08-18

This tweet notes that while the Transformer is not quite a Ship of Theseus, best practices have radically changed since 2017 for positional encodings, layer norms, attention mechanisms, residual streams, and MLP components. It argues that architecture work is not dead, and respecting scaling laws involves more than just scaling up.

Original post →

More from Research

Research channel →