Inference capacity is not a liquid market — book 3 months ahead, not 3 weeks

saranormous · x · 2026-08-21

Sara Hooker argues that once a company scales, inference runs into real compute constraints: the inference capacity market is not liquid, and getting capacity on short notice is hard. Locking in capacity 3 months out is far easier than 3 weeks, yet end users don't care — they just expect the service to work. Her advice: stay on good terms with your inference vendor.

Related event: Compute Crunch Spreads to Inference as GPU Capacity Becomes Illiquid(3 posts)→

Original post →

More from Infra

Infra channel →