Long-context bottleneck is training data and cost, not attention, says model author
antoine_chaffin · x · 2026-09-21
Responding to a debate about his model's long-context claims, author antoinechaffin explains that the main limitations for long context have always been training data and cost. Hybrid local + global attention lowers cost, but natural long-context training data is scarce and all models degrade vastly past a certain length. The model was trained for 8k; with RoPE you can change theta and train on longer data to extend it.
Related event: Model Author Defends Long-Context Claims, Says RoPE Tuning Extends Length(2 posts)→
More from Models
- GLM 5.3 Flash, DeepSeek V4.1 Flash and Qwen 3.8 all beat 'frontier' models from just 10 months ago — burny_tech · 2026-09-21
- Reddit meme: Grok 'devolves' while local MiniMax 3 wins users over — Mystvearn_ · 2026-09-21
- Two years after 'intelligence too cheap to meter', $10/$50 models are the new norm — teortaxesTex · 2026-09-21
- Gary Marcus: "General" LLMs Fall Apart the Moment They Leave Their Training Distribution — GaryMarcus · 2026-09-21
- AI hallucinated a passenger's middle name, almost costing him his flight — menhguin · 2026-09-21
- Distilled Qwen3.8-35B-A3B GGUF quantized model trends on Hugging Face — empero-ai · 2026-09-21