Epoch AI: GPT-5.6 latency scales quadratically with context, Claude 5 stays near-linear
dl_weekly · x · 2026-09-17
An Epoch AI report measuring time-to-first-token up to 1M-token contexts found GPT-5.6 (Terra/Sol) latency scales quadratically with prompt length, while Claude Sonnet 5 and Opus 5 scale near-linearly. This aligns with pricing differences—OpenAI doubles input pricing past 272K tokens while Claude keeps flat pricing—and suggests the two families made very different long-context architectural choices. Three noise-robust estimators confirmed the curvature gap; code and data are public.
More from Models
- Databricks rolls out Astra to all ~3500 engineers, coding spend jumps 60% — pwendell · 2026-09-17
- GoBench: LLMs hit 2500 Elo on 9x9 Go vs KataGo's 4400, r=0.83 with ARC-AGI 2 — Roland31415 · 2026-09-17
- User reports Codex usage wiped to zero after buying a reset despite no usage — seatedro · 2026-09-17
- Union Alpha scores 74 on DeepSWE matching GPT-6 Astra, but benchmark may be saturated — brandon_galang · 2026-09-17
- Can a 2.5B Model Plus Modern Harness Match Pre-March-2025 Frontier Models? — COMPLOGICGADH · 2026-09-17
- GPT-6 Astra's Epoch ECI Score Revised Down on Weak Long-Horizon Software Engineering — Jsevillamol · 2026-09-17