AMD and Cerebras Partner on Disaggregated AI Inference System
Sethwinterroth · x · 2026-07-24
AMD and Cerebras have announced a partnership to develop a disaggregated AI inference system that splits workloads across both platforms. AMD Helios will handle prompts and long context windows, while Cerebras’ Wafer-Scale Engine (WSE) focuses on low-latency token generation.
The companies claim that this combined setup could deliver up to 5x more tokens per second per watt compared to Cerebras alone. Initial availability is planned through Cerebras Cloud in the second half of 2026.
Related event: AMD and Cerebras Partner on Disaggregated Ultra-Low Latency AI Inference(6 posts)→
More from Infra
- Intel Shares Surge 11% as AI Demand Drives Stronger-Than-Expected Earnings — econoar · 2026-07-24
- A 3 GW data-center load drop briefly stressed the PJM grid in Northern Virginia — Annual_Judge_7272 · 2026-07-24
- AMD Helios looks strong, but the Vera Rubin comparison is not apples to apples — karlfreund · 2026-07-24
- Artificial Analysis puts model intelligence and cost on San Francisco billboards — ArtificialAnlys · 2026-07-24
- NVIDIA’s Vera Rubin NVL72 cluster lands with 72 GPUs in one rack-scale system — rohanpaul_ai · 2026-07-24
- Intel’s Q2 2026 results land on the radar for AI infrastructure watchers — BenBajarin · 2026-07-24