New Paper: Compute-optimal Is Not Cluster-optimal

yuxin_tang · hn · 2026-08-14

A new paper argues that traditional "compute-optimal" scaling laws do not equate to "cluster-optimal" outcomes.

The research integrates the systems engineering stage directly into the scaling-law stage, emphasizing that candidate architectures should be evaluated based on what the training cluster can actually deliver. The core finding reveals that the optimal level of sparsity for a Mixture-of-Experts (MoE) model heavily depends on the specific cluster hardware configuration used for training.

Original post →

More from Infra

Infra channel →