Epoch AI: GPT's quadratic long-context latency and pricing tiers hint at different architecture
Epoch AI measurements show GPT-5.6's time-to-first-token grows quadratically with context length, with a price jump beyond 272k input tokens, while Claude 5's pricing differs — suggesting divergent underlying architectures.
2026-09-09 ~ 2026-09-11 · 2 related posts
- Epoch AI: GPT's Quadratic Latency vs Claude's Linear May Explain Pricing Gap — scaling01 · 2026-09-09
- Epoch AI: GPT long-context latency scales quadratically, matching price jumps — Jsevillamol · 2026-09-11