Epoch AI: GPT latency curves bend at long context while Claude stays linear

krishnan · x · 2026-09-16

Epoch AI's September 12 brief measured time-to-first-token from 50,000 to 900,000 input tokens and found that models with similar million-token context windows have very different latency curves.

For anyone building long-running agents on long contexts, this directly affects latency and cost design.

Original post →

More from Models

Models channel →