Long-context bottleneck is training data and cost, not attention, says model author

antoine_chaffin · x · 2026-09-21

Responding to a debate about his model's long-context claims, author antoinechaffin explains that the main limitations for long context have always been training data and cost. Hybrid local + global attention lowers cost, but natural long-context training data is scarce and all models degrade vastly past a certain length. The model was trained for 8k; with RoPE you can change theta and train on longer data to extend it.

Related event: Model Author Defends Long-Context Claims, Says RoPE Tuning Extends Length(2 posts)→

Original post →

More from Models

Models channel →