Marin releases 23T-token pretraining dataset for public download

joecole · x · 2026-08-22

Stanford's Marin model pretraining data is now publicly available, containing 23 trillion tokens, downloadable from an S3 bucket. Dataset composition, architecture, infra, and kernel details are on GitHub, with live training on wandb. The team highlights the joint effort across data, architecture, infra, and kernels.

Original post →

More from Infra

Infra channel →