Epoch AI: GPT-5.6 latency scales quadratically with context, Claude 5 stays near-linear

dl_weekly · x · 2026-09-17

An Epoch AI report measuring time-to-first-token up to 1M-token contexts found GPT-5.6 (Terra/Sol) latency scales quadratically with prompt length, while Claude Sonnet 5 and Opus 5 scale near-linearly. This aligns with pricing differences—OpenAI doubles input pricing past 272K tokens while Claude keeps flat pricing—and suggests the two families made very different long-context architectural choices. Three noise-robust estimators confirmed the curvature gap; code and data are public.

Related event: Epoch Benchmarks: GPT-5.6 Latency Grows Quadratically, Claude Near-Linear at Long Context(2 posts)→

Original post →

More from Models

Models channel →