Teknium says Anthropic still leads on long-context coherence and cost
Teknium · x · 2026-07-22
Teknium says long-context quality and affordability remain the main bottleneck for practical model use.
After testing GPT-5.6 Sol and Terra, the post argues that Anthropic is still the clear leader in long-context coherence and may even be cheaper at that scale. It claims OpenAI charges about 2x once prompts exceed 350K tokens, but the models are not worth using there because coherence falls apart. By contrast, the author says Opus and Fable stay coherent even at 800K+ context.
The post also wonders what happened to Magic.dev’s promised 100M-token context window, suggesting that the idea may not have worked out in practice.
More from Models
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11