Hosting a 3T Model on Cerebras Requires 68 Chips and 1.5MW

zephyr_z9 · x · 2026-08-14

Addressing the issue of attention dominating at longer contexts, the author discusses scaling attention and FFN independently. They speculate that future systems will aggressively use context compaction systems and add more Blackwell chips on the attention side.

Furthermore, a rough cost estimate is provided for deploying a 3-trillion parameter model (like Sol) on Cerebras architecture. Since a single WSE chip has about 44GB of SRAM, hosting this model would require around 68 chips. This would cost nearly $200M and consume around 1.5MW of power for a single copy.

Related event: Cerebras and Blackwell Launch Commercial AFD Architecture(2 posts)→

Original post →

More from Infra

Infra channel →