Epoch AI: GPT's Quadratic Latency vs Claude's Linear May Explain Pricing Gap
scaling01 · x · 2026-09-09
Epoch AI finds OpenAI GPT-5.6 gets pricier past 272k input tokens while Claude 5 pricing stays flat. Their TTFT measurements suggest why: GPT shows a noticeable quadratic latency scaling with context length, while Claude stays near-linear — hinting at underlying architectural differences.
More from Infra
- DeepSeek v4 and GLM Now Run Faster Than vLLM and SGLang — jedisct1 · 2026-09-09
- Estha Turns One Mac Into a Shared Local AI Server for a Whole Team — HaktanSuren · 2026-09-09
- Alexandr Wang backs claim that Scale's own compute lets it subsidize Muse's speed — alexandr_wang · 2026-09-09
- What 100 GW of compute really means: 876 TWh a year and a country-scale power system — shyamalanadkat · 2026-09-09
- vLLM's hard-won lessons: pipeline parallelism falters on warm agent turns, 2.7x decode on Kimi K3 — vllm_project · 2026-09-09
- vLLM details full-stack optimizations for real-world agentic serving on AgentX benchmark — vllm_project · 2026-09-09