Model Author Defends Long-Context Claims, Says RoPE Tuning Extends Length

Model author antoinechaffin defended his model's long-context capability, saying it was trained at 8k and can be extended by adjusting RoPE theta, and that the real bottleneck for long context is scarce training data and cost rather than the hybrid attention design.

2026-09-21 ~ 2026-09-21 · 2 related posts

Full story(2 episodes)→