Llama 3.1 405B Hits 5.6k t/s on Cerebras for Select Customers
kimmonismus · x · 2026-08-01
Cerebras has rolled out its 5.6k tokens/s inference speed, but it is currently limited to select customers. This breakthrough is achieved by running the Llama 3.1 405B model on their Wafer-Scale Engine (WSE).
More from Infra
- antirez looks into deploying LLMs on DGX Spark — antirez · 2026-08-01
- Modal Releases Comprehensive GPU Glossary Covering Hardware to Software Stack — charles_irl · 2026-08-01
- ARM Introduces FEAT_CSSC: Native Popcount for General-Purpose Registers — lemire · 2026-08-01
- Train Your Own Model When Inference Exceeds $750/Day: Pallet's Playbook — marcbhargava · 2026-08-01
- Local Inference of 91GB Audio Model: 127GB RAM Needed for 1M Context — andimarafioti · 2026-08-01
- MediaTek Expects 400G/448G SerDes IP Ready by H2 Next Year — rwang07 · 2026-08-01