1M-token context may be a paper spec: effective window capped by internal attention dimensions
AlexTensor · x · 2026-10-11
Chomba Bupe argues that a model's advertised maximum context window (MCW) of e.g. 1M tokens is not what users actually get: the usable maximum effective context window (MECW) is capped by the model's internal dimension D, in which attention is computed, and is much lower than the spec.
- Core claim: spec-sheet context length and practically usable context are two different things, with the latter constrained by architecture-level dimensions
- In the quoted reply, DenLoginoff asks whether it's fair to say even frontier models with 1M-token windows have a much smaller "effective" window, calling it a deeper analysis of the math underpinnings
Note this is an individual's analysis raising questions about long-context marketing claims, not an official finding.
More from Models
- Claude defaults to Anthropic's signature orange — likely a deliberate choice — deanwball · 2026-10-11
- Gemini 4 Argon reportedly delayed to October 20, unverified rumor — ThePrimeNumbers · 2026-10-11
- Google Now Blocks Taylor Swift Lyrics Searches as AI Safeguards Get Too Blunt — evijit · 2026-10-11
- Stanford says OpenAI cut ChatGPT Edu Pro usage 50% mid-contract after $600/person upfront payment — RexDouglass · 2026-10-11
- Musk explains Cursor acquisition: coding data, talent, compute and enterprise sales were decisive — ns123abc · 2026-10-11
- Gemini 3.8 Flash Thinks of Pausing as "Time Travel" — repligate · 2026-10-11