Stanford Details How Marin 535B Merged 152 Datasets into 25T Tokens
Stanford's Marin team explained how it cleaned and integrated 152 licensed Hugging Face datasets into 25T tokens for training the open-source 535B model, detailing the full data engineering pipeline.
2026-09-24 ~ 2026-09-24 · 2 related posts
- Inside Marin's 535B-token run: integrating 25T tokens from 152 open datasets — jyangballin · 2026-09-24
- Stanford Marin's 535B run: building 25T tokens from 152 open Hugging Face datasets — anshulkundaje · 2026-09-24