Dev explains: model trained at 8k context, RoPE theta tweak extends further
antoine_chaffin · x · 2026-09-21
Model author antoinechaffin answers context-length questions:
- Raising maxlen (README default 512, up to 2048/4096/8192) doesn't hurt long-context performance; 8k is what the model was trained for.
- Architecture detail: the encoder uses 18 sliding-window layers at 128 tokens plus 10 full-attention layers.
- Technically RoPE-based, so you can change theta and train on longer data to extend context.
Related event: Model Author Defends Long-Context Claims, Says RoPE Tuning Extends Length(2 posts)→
More from Models
- Codex computer use tested: Astra works while Luna and Sol fail on Plus plan — FamilyNP · 2026-09-21
- service_tier=fast rejected on ChatGPT subscription, API-only parameter confirmed — TrickyPlastic · 2026-09-21
- Chinese open-source labs explode on OpenRouter: Moonshot +2425%, Z.ai +1925%, DeepSeek +1000% — FinanceYF5 · 2026-09-21
- Yacine: coding LLMs produce 'total complex garbage' — I still read every line — yacineMTB · 2026-09-21
- Researcher: LLMs write convincing related work, but convincing isn't comprehensive — lucacarlone1 · 2026-09-21
- OpenAI's secret technique for upcoming Astra model sparks security concerns — keviv9 · 2026-09-21