Stanford's open Marin 535B-A23B run kicks off: 18.75T tokens, 2.7e24 FLOPs, ~3 months on GB200
burny_tech · x · 2026-09-07
Percy Liang announced that Marin 535B-A23B started training this week, fully in the open: pretraining (80%) plus midtraining (20%) on 18.75T tokens across 11 GB200 NVL72 nodes for 3 months (2.7e24 FLOPs), with post-training to follow. Before launch, the team trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues and forecast the hero run — their biggest yet. Quoting, NicholasBardy notes how aggressive the early-y slope is for larger models while the rest of the curve stays relatively flat: 'feels like we're really missing something.'
Related event: Stanford's Marin Starts Fully Open Training of 535B-A23B Model(2 posts)→
More from Infra
- Interactive speculative decoding tutorial built for NeurIPS Education Track — Madisonkanna · 2026-09-07
- Nvidia guides FY28 to ~$691B, non-hyperscaler demand growing 100% a year — Beth_Kindig · 2026-09-07
- Mike Frank: AI nails reversible computing theory, but the engineering is brutally hard — MikePFrank · 2026-09-07
- Wan2GP lands on Pinokio: one-click AI video generation for 6GB+ VRAM machines — cocktailpeanut · 2026-09-07
- Wan2GP AMD edition hits Pinokio, supporting all RDNA 2-4 discrete GPUs — cocktailpeanut · 2026-09-07
- Buying a room full of hardware to run OpenClaw as supreme rage bait — HankYeomans · 2026-09-07