Hybrid attention matches full attention on long context, 'just cheaper,' says author
antoine_chaffin · x · 2026-09-21
Responding to AIQuanting's observation that 18 of the model's 28 layers use 128-token sliding windows with only 10 full-attention layers, author antoinechaffin says it's not an issue: long-context performance is equivalent to full attention across every layer — 'it's just cheaper.'
More from Models
- Codex computer use tested: Astra works while Luna and Sol fail on Plus plan — FamilyNP · 2026-09-21
- service_tier=fast rejected on ChatGPT subscription, API-only parameter confirmed — TrickyPlastic · 2026-09-21
- Chinese open-source labs explode on OpenRouter: Moonshot +2425%, Z.ai +1925%, DeepSeek +1000% — FinanceYF5 · 2026-09-21
- Yacine: coding LLMs produce 'total complex garbage' — I still read every line — yacineMTB · 2026-09-21
- Researcher: LLMs write convincing related work, but convincing isn't comprehensive — lucacarlone1 · 2026-09-21
- OpenAI's secret technique for upcoming Astra model sparks security concerns — keviv9 · 2026-09-21