Marin starts training 535B-A23B open model on 18.75T tokens with 11 GB200 NVL72s

_ScottCondron · x · 2026-08-23

Percy Liang's Marin lab kicked off training of a 535B-parameter (23B active) model this week, with the whole process open as usual. The voyage plan: pretraining (80%) + midtraining (20%) over 18.75T tokens on 11 GB200 NVL72 systems for 3 months (2.7e24 FLOPs), followed by post-training. Before the hero run, the team trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) up to 27.7B-A1.2B (926B tokens) to debug issues and forecast the main run — their biggest training run yet.

Related event: Stanford's Marin Open-Sources 23T Tokens of Pretraining Data and Launches Fully Transparent 535B Model Training(6 posts)→

Original post →

More from Research

Research channel →