ModernBERT's 8192 context: 18 of 28 layers are actually 128-token sliding windows
AIQuanting · x · 2026-09-21
In a technical exchange over ModernBERT-style long-context claims, AIQuanting notes that while the README recommends raising maxlen to 2048/4096/8192, changing the config doesn't alter the attention structure: 18 of the model's 28 layers use 128-token sliding window attention, with only 10 full-attention layers. So most of the stack reads locally even with 8192 tokens available, complicating the 'supports 8192' framing.
More from Models
- Reddit meme: Grok 'devolves' while local MiniMax 3 wins users over — Mystvearn_ · 2026-09-21
- Two years after 'intelligence too cheap to meter', $10/$50 models are the new norm — teortaxesTex · 2026-09-21
- Gary Marcus: "General" LLMs Fall Apart the Moment They Leave Their Training Distribution — GaryMarcus · 2026-09-21
- AI hallucinated a passenger's middle name, almost costing him his flight — menhguin · 2026-09-21
- Distilled Qwen3.8-35B-A3B GGUF quantized model trends on Hugging Face — empero-ai · 2026-09-21
- Codex users blast OpenAI over opaque usage resets: 'tell us the limit and refill rate' — StewartalsopIII · 2026-09-21