Stanford Marin's 535B run: building 25T tokens from 152 open Hugging Face datasets
anshulkundaje · x · 2026-09-24
- Percy Liang (Stanford) details Marin's open-data work: the 535B run was built on 25T tokens from 152 licensed-permissive datasets on Hugging Face.
- The team built scalable infrastructure to ingest HF's open training data, performing deduplication, decontamination, and mixing to produce the 25T tokens.
- An accompanying thread covers the full pipeline between downloading those datasets and training a model, reflecting Marin's strong belief in the open community.
Related event: Stanford Details How Marin 535B Merged 152 Datasets into 25T Tokens(2 posts)→
More from Infra
- Open-source ComfyUI nodes losslessly compress models: 28GB to 19GB, bit-identical — New-Shift6661 · 2026-09-24
- NVIDIA demos Nemotron 3.5 Lightning running locally on DGX Spark and Station — NVIDIA Developer · 2026-09-24
- For LLM workloads, AMD CCD count matters more than core count: 16-core tops out at 125GB/s — HankYeomans · 2026-09-24
- If 1 billion people ran personal AI agents, CPU and memory demands would be staggering — firstadopter · 2026-09-24
- Side project ports most video generation models from PyTorch to Jax for TPU — ceciletamura · 2026-09-24
- Cloudflare's bot checks called out for wasting agents' time and tokens — msg · 2026-09-24