Marin releases 23T-token pretraining dataset for public download
joecole · x · 2026-08-22
Stanford's Marin model pretraining data is now publicly available, containing 23 trillion tokens, downloadable from an S3 bucket. Dataset composition, architecture, infra, and kernel details are on GitHub, with live training on wandb. The team highlights the joint effort across data, architecture, infra, and kernels.
More from Infra
- LLM Prompting Wastes Computation; Reuse Potential is Huge — miniapeur · 2026-08-22
- SGLang's Weight Cache Daemon Cuts 1T Model Restart Time from 8.8 min to 32 sec — xiaosun86 · 2026-08-22
- Agent recursive loops blow up context costs: 5% failures eat 25% of bill — MaverikSh · 2026-08-22
- llama.cpp ships version 0.2.0 with official release notes — PhilippeEiffel · 2026-08-22
- Open Source Tool Mark Cleaner Locally Removes AI Text Watermarks and Metadata — VraserX · 2026-08-22
- Paper Reveals Larger LLMs Tolerate More Data Repetition During Pretraining — heghbalz · 2026-08-22