Marin 535B training starts with full open process and scaling ladder

ysu_nlp · x · 2026-08-22

Marin 535B-A23B started training this week with the entire process open. The model will train on 11x GB200 NVL72 clusters for 3 months on 18.75T tokens (2.7e24 FLOPs). Before the hero run, a scaling ladder from 1.6B to 27.7B parameters was trained to debug issues and forecast performance.

Related event: Stanford's Marin Open-Sources 23T Tokens of Pretraining Data and Launches Fully Transparent 535B Model Training(6 posts)→

Original post →

More from Infra

Infra channel →