Model performance degrades in long context; token efficiency varies widely across labs
zakelfassi · x · 2026-08-30
The author argues that large context windows (e.g., 1M) alone are no longer impressive, citing significant performance drops as context grows (e.g., GPT-5.6 Sol dropped from 92.4% at 128K to 61.9% at 512K on needle-in-a-haystack). Different model families show wildly different token efficiencies, even per task. This suggests we are still just scratching the surface of understanding what happens under the hood, potentially leading to 2-3 new emerging disciplines.
More from Models
- heretic: fully automatic censorship removal for LLMs nears 29k stars — p-e-w · 2026-08-30
- Experiment: Claude Easily Assisted in Piracy and Reverse Engineering via agents.md — adonis_singh · 2026-08-30
- OpenAI dominates browser use while Claude's strength is mostly coding, exec says — bindureddy · 2026-08-30
- Claude Opus 5 Backlash: Benchmarks Soar But Daily Use Fails — gerardsans · 2026-08-30
- 'The curve of the letter b is invisible to the model' — tokenizer meme resurfaces — rickasaurus · 2026-08-30
- Humor Benchmark: Gemini 3.7 Wins, GPT-4o Struggles to Be Funny — scaling01 · 2026-08-30