Studies Reveal Transformer Architecture Choices Hinder Long Context Extension
yoavartzi · x · 2026-08-18
The post highlights two COLM papers exploring transformer architecture ingredients. One shows that conventional choices like qk-norm, GQA, and SWA negatively impact long-context extension, with combinations causing up to 47% performance drops. Another argues the output projection layer creates a gradient bottleneck slowing pretraining. These findings emphasize the need for co-designing architectures, especially under compression constraints.
Related event: COLM papers reveal common Transformer designs hurt long-context scaling(4 posts)→
More from Research
- Converting GMMs ↔ PEFs for fast KLD approximation — FrnkNlsn · 2026-08-24
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24
- New Architecture RHEA: Train 1B Model on 8GB VRAM — zemondza · 2026-08-24
- Trained two 16M-param models to do generative CAD with real physics — debreuil · 2026-08-24
- Claude model helps discover complex structure on S^6, solving 60-year-old math problem — Singularitarian · 2026-08-24