Open 1T-token pretraining corpus released at ~$11 per billion tokens trained
jon_durbin · x · 2026-09-11
Developer jondurbin published the full dataset behind his pretraining run, chutesai/mesh-1t-prod-mix: a trillion-ish-token multi-phase pretraining corpus (no SFT/RL, pretraining only), gated with auto-approval on Hugging Face.
- Stored as flat llama3 uint32 token bins (EOS 128001 separators) with val bins, per-phase mixmanifest. / poolconfig. provenance, and SHA256SUMS
- Phase budgets include 600B tokens (ClimbMix+FinePDFs-Edu, UFW, Stack-code, CC-Math, etc.) and a 285B second phase
- Training cost ran $11 per billion tokens, which he calls insane; MFU also impressively high
- Uses a few nodes from lium.io; a run dashboard is available
More from Infra
- Training a 6-Expert MoE GPT-2 From Scratch on a Single RTX 3090 in 8 Days — rasbt · 2026-09-11
- B3IQ Sells Eight Figures of GPUs in Two Weeks, Bets AI Infra Is a $100B Market — templecrash · 2026-09-11
- It Cost $100 in API Credits for an AI Agent to Install Free Software — MartinGTobias · 2026-09-11
- bartowski unveils per-tensor layout maps for GGUF quantization, tests show across-the-board gains — noneabove1182 · 2026-09-11
- antirez runs DeepSeek v4.1 Flash locally on a 128GB M5 Max, SSD streaming surprisingly fast — antirez · 2026-09-11
- OpenAI could 7x its training compute tomorrow: why open-source models still trail by one generation — soumitrashukla9 · 2026-09-11