Studies Reveal Transformer Architecture Choices Hinder Long Context Extension

yoavartzi · x · 2026-08-18

The post highlights two COLM papers exploring transformer architecture ingredients. One shows that conventional choices like qk-norm, GQA, and SWA negatively impact long-context extension, with combinations causing up to 47% performance drops. Another argues the output projection layer creates a gradient bottleneck slowing pretraining. These findings emphasize the need for co-designing architectures, especially under compression constraints.

Related event: COLM papers reveal common Transformer designs hurt long-context scaling(4 posts)→

Original post →

More from Research

Research channel →