Gimlet Cloud partners with Cerebras for 100MW wafer-scale inference, up to 3,000 tokens/sec

Sethwinterroth · x · 2026-09-28

Gimlet Labs announced a strategic partnership with Cerebras to bring wafer-scale compute to Gimlet Cloud, with 100MW of inference capacity planned and the first datacenter online later this year at speeds up to 3,000 tokens/sec.

The pitch: agent latency compounds across model calls — a 10-minute multi-step task at 100 tokens/sec takes 20 seconds at 3,000. Gimlet is also a launch partner for Cerebras' next-gen CS-4, expected in 2027.

Original post →

More from Infra

Infra channel →