Marin Uses Scaling Laws to Predict Training Trajectories, Kicks Off 535B MoE

dlwh · x · 2026-08-21

Marin's team found that scaling laws can simulate the entire training trajectory. Their recent 67B MoE run hit the loss target within 1%, and arbitrary points during training were also predictable within 1%. Based on this, they used a 'scaling ladder' to simulate the 535B MoE run launched yesterday to catch bugs early. The team emphasized that getting a run off the ground involves herculean efforts beyond architecture, such as curating trillions of tokens and adapting hardware.

Original post →

More from Infra

Infra channel →