Four architecture choices can cost up to 47% of long-context performance; new study releases OlmPool

dair_ai · x · 2026-08-13

New research from Ai2, CMU, and UW finds that four seemingly harmless architecture decisions—normalization, GQA, pretraining context length, and sliding window attention—can together degrade long-context performance by up to 47%. These choices are invisible in short-context loss or validation sets. The team releases OlmPool, 26 comparable 7B models with checkpoints before/after extension, trained over 170,000 GPU hours.

Original post →

More from Research

Research channel →