Stanford's open Marin 535B-A23B run kicks off: 18.75T tokens, 2.7e24 FLOPs, ~3 months on GB200

burny_tech · x · 2026-09-07

Percy Liang announced that Marin 535B-A23B started training this week, fully in the open: pretraining (80%) plus midtraining (20%) on 18.75T tokens across 11 GB200 NVL72 nodes for 3 months (2.7e24 FLOPs), with post-training to follow. Before launch, the team trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues and forecast the hero run — their biggest yet. Quoting, NicholasBardy notes how aggressive the early-y slope is for larger models while the rest of the curve stays relatively flat: 'feels like we're really missing something.'

Related event: Stanford's Marin Starts Fully Open Training of 535B-A23B Model(2 posts)→

Original post →

More from Infra

Infra channel →