Gimlet and Cerebras plan 100 MW capacity, targeting 3,000 tokens/sec inference
Sethwinterroth · x · 2026-09-29
Gimlet Labs is deepening its partnership with Cerebras: the two already serve inference traffic together and are now planning 100 MW of capacity, with the first Cerebras-powered datacenter expected this year and speeds up to 3,000 tokens/sec. Gimlet's pitch is inference disaggregation—mapping each phase of inference to the silicon best suited to it—expanding the throughput and interactivity frontier through a single API, enabling products built around agents finishing in seconds.
Related event: Gimlet Partners with Cerebras on 100MW Inference Cloud(2 posts)→
More from Infra
- Macrocosmos launches iota SDK and Liquid Compute to train on disaggregated global compute — markjeffrey · 2026-09-29
- "Got into datacenters for crypto, making 10000x more in AI" — industry quip — wordgrammer · 2026-09-29
- Starship launch just added ~1% to global internet bandwidth, investors say world isn't pricing it in — juanbenet · 2026-09-29
- GPU shortage: B700 unavailable, RTX 5090 listings hit $10,000 — Dismal-Effect-1914 · 2026-09-29
- Disaggregated Quantization Boosts 1-bit LLM Accuracy by 32+ Points and TTFT by 1.78x — ISTA-DASLab · 2026-09-29
- New GPU Prices API Tracks Real-Time H100/B200 Rental Rates via REST or MCP — virattt · 2026-09-29