Open-source 535B MoE model Marin begins training

stanfordnlp · x · 2026-08-22

Stanford NLP started training the open-source Marin 535B-A23B model. It will run on 11 x GB200 NVL72 clusters for 3 months on 18.75T tokens (2.7e24 FLOPs), split into 80% pretraining and 20% midtraining. The team debugged via a scaling ladder from 1.6B to 27.7B. The entire process is open to promote transparency.

Related event: Stanford's Marin Launches Fully Transparent Training of 535B Model, Open-Sources 23T Token Dataset(7 posts)→

Original post →

More from Infra

Infra channel →