a16z backs Gimlet Labs' multi-silicon inference cloud promising 10X throughput in same power envelope
a16z Newsletter · rss · 2026-09-05
a16z announced its investment in Gimlet Labs, calling it the first multi-silicon inference cloud designed to "produce more intelligence from every watt."
Why now
AI inference is one of the fastest-growing markets in history, yet every physical input is scarce: powered land, turbines, transformers, data centers, GPUs, advanced-node wafers, HBM. GPU leasing is at record prices; OpenAI and Anthropic are constrained by compute expansion. Five US hyperscalers are expected to spend $1 trillion in capex next year, and NVIDIA has mobilized $500B+ for AI factories.
Inference is not one workload
Voice assistants need latency, batch needs throughput, agents juggle small specialized and large reasoning models, tool calls, and CPU execution — and a single model call splits into compute-bound prefill and memory-bound decode. No single processor wins everywhere; GPUs will increasingly work alongside CPUs and purpose-built silicon. The hard part is making heterogeneous hardware operate as one system (even inlet-water temperatures differ between Cerebras and NVIDIA rigs).
Gimlet's approach
Gimlet builds an execution plan per workload balancing latency, throughput, and cost: routing models and tools to different processors, splitting prefill/decode, or dividing a model at the layer/operation level, with a compiler and runtime coordinating it all behind a single inference API. It also pools GPUs, CPUs, and accelerators at the physical data center level, managing differing networking, power, and cooling. It claims up to 10X gains in throughput and interactivity for frontier models within the same power envelope, with a frontier lab and a hyperscaler already as customers.
More from Venture
- Cathie Wood: Block's product velocity is phenomenal since Jack reorganized around AI — CathieDWood · 2026-09-05
- Rebutting Lessin: enterprise AI is servicizing products — we're getting digital teammates — arieljalali · 2026-09-05
- ChapterPal hits 6,160 users in year one, auto-adds 2 books and 10 papers daily — burkov · 2026-09-05
- Why 1% startup equity is often the worst zone, by the numbers — isaacinthesky · 2026-09-05
- $2 Trillion Wiped From SaaS, Then a Highly Uneven Rebound: Winners and Losers of 2026 — AccBalanced · 2026-09-05
- Investors of Harvey and Legora trade barbs over growth rates in legal AI — vaibhavbetter · 2026-09-05