Cerebras unveils CS-4, nearly doubles AI inference performance without new process node
kimmonismus · x · 2026-08-20
Cerebras unveiled the CS-4 system, nearly doubling its AI inference performance without moving to a new process node. By redesigning power delivery and cooling, the wafer runs at double the clock speed while retaining the same 5nm process, 4 trillion transistors, and 900,000 AI cores.
Key Specs per WSE-3 Turbo:
- 250 PFLOPs of AI compute
- 43.2 PB/s of memory bandwidth
- 2.4 Tb/s of I/O bandwidth
A single CS-4 rack combines three wafers for 750 PFLOPs and 129.6 PB/s of memory bandwidth. On GPT-OSS-120B, Cerebras reports over 4,400 tokens per second per user, making it up to 30x faster than GPU-based systems.
Related event: Cerebras Unveils CS-4 Wafer-Scale AI Inference System(16 posts)→
More from Infra
- Perplexity Launches Portable Computer, a Local-First Agent Stack on DGX Spark — ChrisUniverse · 2026-08-27
- RootCrak builds x402 security layer for autonomous agent transactions — Thionne_WTZ · 2026-08-27
- Max Hodak: Anonymous model testing routed data to Chinese datacenter — ohlennart · 2026-08-27
- Opinion: Why Targeting Data Centers is an Environmentalist Mistake — AndyMasley · 2026-08-27
- Advocating for Independent Secure Clusters: Open Science Needs Open Compute — gajesh · 2026-08-27
- Self-hosting LLMs on Budget Hardware: Principles, Optimization, and Benchmarks — jflesch · 2026-08-27