Compute Will Dominate the Frontier Model Era
deedydas · x · 2026-07-20
The author argues that the main theme of AI in the coming years will pivot to compute, driven by rapidly shifting supply and demand for frontier models.
- Taking Kimi K3 as an example: within 2 days of launch, it surged to the top 10 on OpenRouter, processing about 140B tokens daily. However, its inference infrastructure is already strained: throughput dropped from 30 tok/s to 13 tok/s, end-to-end latency spiked to 72 seconds, and time-to-first-token exceeded 20 seconds.
- To adequately equip the "quantized Kimi K3" with compute, the minimum requirement is purchasing 8 B300 GPUs, costing around $500,000. Upgrading to the highly recommended GB300 NVL72 rack would cost about $4 million.
- Core judgment: Moonshot might lack sufficient compute to meet demand. Even US-based inference providers may not expand capacity quickly, as 2.8T-scale models are simply too heavy.
- Many GPU providers and neoclouds have started signing 3- to 5-year long-term contracts with 30% upfront payments, driving compute prices even higher. Top labs, hyperscalers, and some major LLM companies have already locked in resources, leaving remaining players scrambling for scarce capacity.
- Conclusion: Although model unit prices are declining long-term, the cost of "frontier capabilities" hasn't significantly collapsed. With sustained demand growth, continuous capability improvements, and constrained compute supply, companies that lock in compute early will capture greater value.
Related event: Compute to Dominate Frontier AI Era, but with Variables(2 posts)→
More from Infra
- oMLX 0.5.2 adds Mac menu-bar stats, low-bit decode kernels, and faster downloads — awnihannun · 2026-07-22
- Strangeworks launches Aura to turn enterprise ops into production optimization systems — whurley · 2026-07-22
- Graph workload 854.graph500 enters SPEC CPU 2026 as a new CPU benchmark — Prof_DavidBader · 2026-07-22
- HilbertRaum open-sources a fully local AI chat and document analysis app for private use — Vladowski · 2026-07-22
- Hybrid and local inference are emerging as a response to AI energy and token costs — dmitry140 · 2026-07-22
- NVIDIA details Vera CPU with 2x performance claims and a 22,000-core rack — ryanshrout · 2026-07-22