Debate over ModernBERT long-context: sliding-window layers and 8192 limit questioned
On September 21, community members debated the long-context capability of a ModernBERT-style model. AIQuanting pointed out that although the README claims maxlen can be raised to 2048, 4096, or even 8192, longer configurations don't change the attention architecture itself: of the 28 layers in the config, 18 use local sliding-window attention over 128 tokens, with only 10 layers doing full-sequence attention. He also questioned the measured setup: maxpositionembeddings is set to 8192, but the model card (context row) lists 512, and the agent config defaults maxlen to 512—a gap between marketing and actual usage.
Confirmed
- 18 of the model's 28 layers use 128-token sliding-window attention; 10 layers use full-sequence attention (per AIQuanting's reading of the config files).
- maxpositionembeddings is configured as 8192, while the model card lists 512 and the agent default maxlen is 512.
- Model author antoinechaffin responded that the hybrid sliding-window + global-attention design matches per-layer full attention on long-context performance, only at lower compute cost.
- JFPuget explained that 8192 is the model's supported upper bound; if downstream uses (e.g., the laya card, agent configs) impose smaller limits, that's a downstream choice, not a model capability issue.
Why it matters
- Context length is a key metric for model selection, but inconsistencies among config fields, model cards, and downstream defaults can mislead users, who need to distinguish "architecture-supported limits" from "effective length in practice."
- Whether sliding-window + global hybrid attention truly matches full-attention performance on long-context tasks is the core controversy for cheap long-context approaches; the author's "same performance, just cheaper" response hasn't settled community disagreement over the measured setup.
2026-09-21 ~ 2026-09-21 · 5 related posts
- Episode 1: Debate over ModernBERT long-context: sliding-window layers and 8192 limit questioned(2026-09-21, 5 posts)
- Episode 2: Model Author Defends Long-Context Claims, Says RoPE Tuning Extends Length(2026-09-21, 2 posts)
Primary sources
- ModernBERT's 8192 context: 18 of 28 layers are actually 128-token sliding windows — AIQuanting ·
- Hybrid attention matches full attention on long context, 'just cheaper,' says author — antoine_chaffin ·
- ModernBERT supports 8192 tokens natively — downstream 512 limits can be safely removed — JFPuget ·
- ModernBERT's 28 layers aren't uniform: 18 slide a 128-token window, only 10 attend fully — AIQuanting · 2026-09-21
- [source] Hybrid attention matches full attention on long context, 'just cheaper,' says author — antoine_chaffin · 2026-09-21
- Config says 8192, model card says 512: which run produced the long-context number? — AIQuanting · 2026-09-21
- [source] ModernBERT supports 8192 tokens natively — downstream 512 limits can be safely removed — JFPuget · 2026-09-21
- [source] ModernBERT's 8192 context: 18 of 28 layers are actually 128-token sliding windows — AIQuanting · 2026-09-21