Epoch AI: GPT's Quadratic Latency vs Claude's Linear May Explain Pricing Gap

scaling01 · x · 2026-09-09

Epoch AI finds OpenAI GPT-5.6 gets pricier past 272k input tokens while Claude 5 pricing stays flat. Their TTFT measurements suggest why: GPT shows a noticeable quadratic latency scaling with context length, while Claude stays near-linear — hinting at underlying architectural differences.

Original post →

More from Infra

Infra channel →