Inside Marin's 535B-token run: integrating 25T tokens from 152 open datasets
jyangballin · x · 2026-09-24
Marin team member William Barr Held published a long thread on the engineering between downloading open training data and actually training a model.
- Marin's 535B-token run built on 25T tokens from 152 datasets with licenses permitting training
- Covers the full pipeline: downloading, cleaning, deduplicating, and unifying massive open datasets from Hugging Face
- A systematic reference for teams building open pretraining data pipelines
More from Research
- Two months of π0.5 finetuning ablations: diversity beats raw scaling, 98% success — philfung · 2026-09-24
- Researcher: AI's autonomous bioinformatics work still short of real discovery — dvir_a · 2026-09-24
- Beyond Repeated Sampling: Scaling LLM Test-Time Reasoning With Learned Concepts — rbhar90 · 2026-09-24
- CoRL 2026 to Host Workshop on Learning from Corrections and Interventions, CFP Open — ebiyik_ · 2026-09-24
- Anthropic's Enzyme Discovery Is Cool but Very Preliminary, Researchers Caution — iScienceLuvr · 2026-09-24
- RL tutorial: using Jev as reward model lifts Qwen reward from 0.583 to 0.759 — sophiamyang · 2026-09-24