Open-source Marin kicks off 535B/23B MoE pretraining run, fully documented in public
dlwh · x · 2026-09-03
Open Athena's Marin project launched its capstone run: a 535B-total/23B-active MoE model pretraining on 18T tokens (agentic, coding, science data) for 3 months, fully documented on GitHub. The run is tracking predictions closely. Past year highlights: Delphi scaling suite predicting 300x-larger models within 0.2%, Marin MoE V1 (129B/16B) landing within 1% of preregistered loss with 6.7x theoretical (3.6x realized on TPU v4) speedup over dense, the Iris global compute scheduler, DataKit data curation, and the upcoming Snowball 67B-A2B model with 262K context. Team grew from 1 FTE to 10.
Related event: Open Athena Launches Marin, a 535B-Parameter MoE Trained in the Open(2 posts)→
More from Infra
- E2B sandbox runs RL rollouts up to 3x faster, cutting idle GPU time and training cost — badphilosopher · 2026-09-04
- ChatGPT, Claude, Codex and Grok go down in the same window; OpenAI and Anthropic confirm elevated errors — rohanpaul_ai · 2026-09-04
- "Global productivity just dropped to zero": meme captures the mass AI outage across xAI, OpenAI, Anthropic, Google and AWS — TheZachMueller · 2026-09-04
- Uzu lab ships speculative decoding implementation, launching with Qwen3.6 27B support — awnihannun · 2026-09-04
- NVIDIA at IFA 2026: 1.9x faster local inference, PAIR router, RTX Spark PCs in October — nordicinst · 2026-09-04
- VPIPE is also one of the fastest local image generators: Krea 2 in ~45s on M5 Pro — TgoAI · 2026-09-04