a16z Essay: AI Inference Demand to Drive Shift from General GPUs to Custom Silicon
a16z Newsletter · rss · 2026-07-23
A deep dive essay from a16z argues that AI inference is rapidly overtaking model training as the largest workload in computing history, driving a historic shift from general-purpose GPUs to custom specialized hardware.
Key takeaways include:
- Inference as the Major Cost: While training is a one-time capital expenditure, inference is an operating expense that scales linearly with usage. Generating a single token requires massive memory reads and matrix arithmetic. With the explosion of users and agent platforms at OpenAI and Anthropic, the demand and cost for inference compute are going vertical.
- Bottlenecks of General-Purpose GPUs: GPUs launched the AI era due to their flexibility, but they are fundamentally mismatched for the highly specific workload of inference. To maximize throughput, the industry relies on "batching," which inherently increases latency—a tradeoff unacceptable for real-time coding agents. Furthermore, as data centers become power-constrained, watts wasted on a GPU's general-purpose flexibility directly cut into token-generating revenue.
- The Rise of Custom Silicon: Just as graphics processing migrated from CPUs to GPUs, massive and stable workloads inevitably earn their own specialized hardware. Google's TPU is an early example, but the trend is now accelerating across hyperscalers and frontier labs.
- Etched's Bet: The essay highlights the startup Etched, whose founders dropped out of Harvard before ChatGPT to build an end-to-end inference system—from chips to racks. Their core technologies include low-voltage inference (packing more compute into the same power envelope) and cluster-scale memory (pooling memory across an entire rack via ultra-low latency interconnects). Etched has hired over 400 engineers, operating on the philosophy that "production is the product."
More from Venture
- AMD: AI Inference Compute Surpasses Training, Targeting $2T TAM by 2030 — ryanshrout · 2026-07-24
- Greylock says consumer AI can escape OpenAI’s “God Box” through vertical services and networks — SethGRosenberg · 2026-07-24
- Paper raises $34 million Series A from Accel and ICONIQ Capital — hewarsaber · 2026-07-24
- Gary Marcus says AI-heavy tech stocks now trade like crypto — GaryMarcus · 2026-07-24
- Multi-agent tools that share workspace and approvals cut task cost to $1.65 — HeyNayeem · 2026-07-23
- Cognition is buying Interaction, the team behind the Poke product — marvinvonhagen · 2026-07-23