Model Author Defends Long-Context Claims, Says RoPE Tuning Extends Length
Model author antoinechaffin defended his model's long-context capability, saying it was trained at 8k and can be extended by adjusting RoPE theta, and that the real bottleneck for long context is scarce training data and cost rather than the hybrid attention design.
2026-09-21 ~ 2026-09-21 · 2 related posts
- Episode 1: Debate over ModernBERT long-context: sliding-window layers and 8192 limit questioned(2026-09-21, 5 posts)
- Episode 2: Model Author Defends Long-Context Claims, Says RoPE Tuning Extends Length(2026-09-21, 2 posts)
- Dev explains: model trained at 8k context, RoPE theta tweak extends further — antoine_chaffin · 2026-09-21
- Long-context bottleneck is training data and cost, not attention, says model author — antoine_chaffin · 2026-09-21