1M-token context may be a paper spec: effective window capped by internal attention dimensions

AlexTensor · x · 2026-10-11

Chomba Bupe argues that a model's advertised maximum context window (MCW) of e.g. 1M tokens is not what users actually get: the usable maximum effective context window (MECW) is capped by the model's internal dimension D, in which attention is computed, and is much lower than the spec.

Note this is an individual's analysis raising questions about long-context marketing claims, not an official finding.

Original post →

More from Models

Models channel →