Open-source 535B MoE model Marin begins training
stanfordnlp · x · 2026-08-22
Stanford NLP started training the open-source Marin 535B-A23B model. It will run on 11 x GB200 NVL72 clusters for 3 months on 18.75T tokens (2.7e24 FLOPs), split into 80% pretraining and 20% midtraining. The team debugged via a scaling ladder from 1.6B to 27.7B. The entire process is open to promote transparency.
More from Infra
- Perplexity goes local: private tasks hand off to on-device small models — HowDevelop · 2026-09-07
- DeepSeek V4 Flash at 75% off via Merge Gateway: $0.04/M input tokens through Sept 30 — shensi · 2026-09-07
- Qwen 3.8 Flash Next runs at 65 t/s on M3 Ultra, Q2 weights released on Hugging Face — ivanfioravanti · 2026-09-07
- The Myth of Self-Hosted AI: 'Local' Models Still Route Through the Cloud — nomad-nostalgia · 2026-09-07
- Homelab With 4x RTX 4090 Weighs vLLM+P2P Patch vs llama.cpp for Qwen Models — dowitex · 2026-09-07
- Could AI run entirely on your phone? It could upend OpenAI's pricing — kevinsurace · 2026-09-07