Transformer Best Practices Radically Changed Since 2017
ChrisGPotts · x · 2026-08-18
This tweet notes that while the Transformer is not quite a Ship of Theseus, best practices have radically changed since 2017 for positional encodings, layer norms, attention mechanisms, residual streams, and MLP components. It argues that architecture work is not dead, and respecting scaling laws involves more than just scaling up.
More from Research
- Prof. teaches AI architectures via hand calculation: Transformer to Mamba — techNmak · 2026-08-18
- Video Model Evaluation Guide: Quantifying Generation Quality — Majumdar_Ani · 2026-08-18
- Study: LLMs as synthetic survey respondents are plausible but not valid — Mantas Lukauskas · 2026-08-18
- Adaption AI Launches Custom Evals for Pro Users — sarahookr · 2026-08-18
- Interactive Diagram: Understanding Autoencoders by Hand — ProfTomYeh · 2026-08-18
- PyLate Merges Back Into Sentence Transformers — mrdrozdov · 2026-08-18