Llama 3.1 405B Hits 5.6k t/s on Cerebras for Select Customers

kimmonismus · x · 2026-08-01

Cerebras has rolled out its 5.6k tokens/s inference speed, but it is currently limited to select customers. This breakthrough is achieved by running the Llama 3.1 405B model on their Wafer-Scale Engine (WSE).

Original post →

More from Infra

Infra channel →