Four architecture choices can cost up to 47% of long-context performance; new study releases OlmPool
dair_ai · x · 2026-08-13
New research from Ai2, CMU, and UW finds that four seemingly harmless architecture decisions—normalization, GQA, pretraining context length, and sliding window attention—can together degrade long-context performance by up to 47%. These choices are invisible in short-context loss or validation sets. The team releases OlmPool, 26 comparable 7B models with checkpoints before/after extension, trained over 170,000 GPU hours.
More from Research
- Deep Dive: The Model Eats the Harness in an Agentic World — pzakin · 2026-08-13
- PyTorch Devs Release Interactive Pipeline Parallelism Scheduling Tutor — ezyang · 2026-08-13
- Sekai2: A 2,800+ Hour Interactive World Modeling Video Dataset — udmrzn · 2026-08-13
- Will Pre-AI Human Data Become More Valuable as the Internet Fills with AI Content? — ArcanuMELO · 2026-08-13
- Modeling Uncertainty in Code Review Agents: Non-Exclusive vs. Mutually Exclusive Risks — Accomplished-Fun4629 · 2026-08-13
- Embodied AI Breakthrough: SONIC System for Robot Motion Tracking Published in Science Robotics — zhengyiluo · 2026-08-13