Model performance degrades in long context; token efficiency varies widely across labs

zakelfassi · x · 2026-08-30

The author argues that large context windows (e.g., 1M) alone are no longer impressive, citing significant performance drops as context grows (e.g., GPT-5.6 Sol dropped from 92.4% at 128K to 61.9% at 512K on needle-in-a-haystack). Different model families show wildly different token efficiencies, even per task. This suggests we are still just scratching the surface of understanding what happens under the hood, potentially leading to 2-3 new emerging disciplines.

Original post →

More from Models

Models channel →