New Paper: Compute-optimal Is Not Cluster-optimal
yuxin_tang · hn · 2026-08-14
A new paper argues that traditional "compute-optimal" scaling laws do not equate to "cluster-optimal" outcomes.
The research integrates the systems engineering stage directly into the scaling-law stage, emphasizing that candidate architectures should be evaluated based on what the training cluster can actually deliver. The core finding reveals that the optimal level of sparsity for a Mixture-of-Experts (MoE) model heavily depends on the specific cluster hardware configuration used for training.
More from Infra
- Tobi Lütke: local Dell server runs DeepSeek 4.1 Flash at ~300 tok/s, a billion tokens a month — BLUECOW009 · 2026-09-21
- Running Qwen3.8-27B EXL3 on RTX 3060 + 5060 Ti: 50 tok/s with tensor parallelism and MTP — bring_back_the_v10s · 2026-09-21
- Baseten CEO says token volume grew 40x YoY while revenue grew ~10x in 12 months — rohanpaul_ai · 2026-09-21
- AI doesn't live in the cloud: who pays the environmental price of scale? — SuzannahB1001 · 2026-09-21
- AI Cluster Bottleneck Isn't Chips — It's the Lasers Moving Data Between Them — McDonaghMatthew · 2026-09-21
- Inside SemiAnalysis: the research firm guiding a $1T AI infrastructure buildout — AccBalanced · 2026-09-21