ModernBERT's 8192 context: 18 of 28 layers are actually 128-token sliding windows

AIQuanting · x · 2026-09-21

In a technical exchange over ModernBERT-style long-context claims, AIQuanting notes that while the README recommends raising maxlen to 2048/4096/8192, changing the config doesn't alter the attention structure: 18 of the model's 28 layers use 128-token sliding window attention, with only 10 full-attention layers. So most of the stack reads locally even with 8192 tokens available, complicating the 'supports 8192' framing.

Related event: Debate over ModernBERT long-context: sliding-window layers and 8192 limit questioned(5 posts)→

Original post →

More from Models

Models channel →