Marin starts training 535B-A23B open model on 18.75T tokens with 11 GB200 NVL72s
_ScottCondron · x · 2026-08-23
Percy Liang's Marin lab kicked off training of a 535B-parameter (23B active) model this week, with the whole process open as usual. The voyage plan: pretraining (80%) + midtraining (20%) over 18.75T tokens on 11 GB200 NVL72 systems for 3 months (2.7e24 FLOPs), followed by post-training. Before the hero run, the team trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) up to 27.7B-A1.2B (926B tokens) to debug issues and forecast the main run — their biggest training run yet.
More from Research
- Pixel32Bench compares language models via 32x32 pixel generation — TheMoonMidas · 2026-08-23
- AI in material science: Opportunities with cloud labs and simulation — JacquesThibs · 2026-08-23
- Marin 535B Training Starts with Full Transparency on FLOPs and Configs — _ScottCondron · 2026-08-23
- Developer calls MCP research paper error-riddled and partially AI-generated — benfielding · 2026-08-23
- NanoGPT Speedrun Leaderboard Summary — RichmanRonald · 2026-08-23
- Where Are All the Prompt Injection Damages? — joshua_saxe · 2026-08-23