AMD and Cerebras Partner on Disaggregated AI Inference System

Sethwinterroth · x · 2026-07-24

AMD and Cerebras have announced a partnership to develop a disaggregated AI inference system that splits workloads across both platforms. AMD Helios will handle prompts and long context windows, while Cerebras’ Wafer-Scale Engine (WSE) focuses on low-latency token generation.

The companies claim that this combined setup could deliver up to 5x more tokens per second per watt compared to Cerebras alone. Initial availability is planned through Cerebras Cloud in the second half of 2026.

Related event: AMD and Cerebras Partner on Disaggregated Ultra-Low Latency AI Inference(6 posts)→

Original post →

More from Infra

Infra channel →