Open-source Marin kicks off 535B/23B MoE pretraining run, fully documented in public

dlwh · x · 2026-09-03

Open Athena's Marin project launched its capstone run: a 535B-total/23B-active MoE model pretraining on 18T tokens (agentic, coding, science data) for 3 months, fully documented on GitHub. The run is tracking predictions closely. Past year highlights: Delphi scaling suite predicting 300x-larger models within 0.2%, Marin MoE V1 (129B/16B) landing within 1% of preregistered loss with 6.7x theoretical (3.6x realized on TPU v4) speedup over dense, the Iris global compute scheduler, DataKit data curation, and the upcoming Snowball 67B-A2B model with 262K context. Team grew from 1 FTE to 10.

Related event: Open Athena Launches Marin, a 535B-Parameter MoE Trained in the Open(2 posts)→

Original post →

More from Infra

Infra channel →