Epoch AI: GPT long-context latency scales quadratically, matching price jumps

Jsevillamol · x · 2026-09-11

Epoch AI measured time-to-first-token (TTFT) scaling and found OpenAI's GPT-5.6 and Anthropic's Claude 5 behave very differently: GPT shows a noticeable quadratic component as context grows while Claude stays closer to linear — mirroring their pricing (GPT prices jump past 272k input tokens; Claude stays fixed).

OpenAI's newly released GPT-6 Astra also raises API pricing beyond 272k tokens, and Epoch's additional latency measurements show similar curvature to GPT-5.6, suggesting an architectural difference rather than pure pricing strategy.

Related event: Epoch AI: GPT's quadratic long-context latency and pricing tiers hint at different architecture(2 posts)→

Original post →

More from Infra

Infra channel →