Epoch AI: GPT's quadratic long-context latency and pricing tiers hint at different architecture

Epoch AI measurements show GPT-5.6's time-to-first-token grows quadratically with context length, with a price jump beyond 272k input tokens, while Claude 5's pricing differs — suggesting divergent underlying architectures.

2026-09-09 ~ 2026-09-11 · 2 related posts