Hybrid attention matches full attention on long context, 'just cheaper,' says author

antoine_chaffin · x · 2026-09-21

Responding to AIQuanting's observation that 18 of the model's 28 layers use 128-token sliding windows with only 10 full-attention layers, author antoinechaffin says it's not an issue: long-context performance is equivalent to full attention across every layer — 'it's just cheaper.'

Related event: Debate over ModernBERT long-context: sliding-window layers and 8192 limit questioned(5 posts)→

Original post →

More from Models

Models channel →