Epoch AI: GPT long-context latency scales quadratically, matching price jumps
Jsevillamol · x · 2026-09-11
Epoch AI measured time-to-first-token (TTFT) scaling and found OpenAI's GPT-5.6 and Anthropic's Claude 5 behave very differently: GPT shows a noticeable quadratic component as context grows while Claude stays closer to linear — mirroring their pricing (GPT prices jump past 272k input tokens; Claude stays fixed).
OpenAI's newly released GPT-6 Astra also raises API pricing beyond 272k tokens, and Epoch's additional latency measurements show similar curvature to GPT-5.6, suggesting an architectural difference rather than pure pricing strategy.
More from Infra
- 2.78T-param Kimi K3 runs inference on a single CPU in 8.24 GB of RAM — udmrzn · 2026-09-11
- Keep the Claude Desktop Workflow, Swap in Local Models via Ollama for Privacy — Technovangelist · 2026-09-11
- LithosAI ships Day-0 API inference for DeepSeek-V4.1-Flash at 250+ tokens/s per user — JiaZhihao · 2026-09-11
- TwelveLabs Marengo 3.0 Goes GA in Amazon Bedrock for Video Semantic Search — AWS ML Blog · 2026-09-11
- 1:26 continuous aerial AI video made entirely on a Mac with MiniMax H3 — cocktailpeanut · 2026-09-11
- KV cache gets QAT too: why this model beats others at fp4 KV cache — stochasticchasm · 2026-09-11