The Stack v3 releases 5T code tokens across 700+ programming languages
LoubnaBenAllal1 · x · 2026-07-23
The Stack team says it is releasing The Stack v3 at a time when strong open models are needed more than ever.
- The dataset contains 5T tokens of code.
- It spans 700+ programming languages.
- The release is positioned as infrastructure for building and training stronger open code models.
Related event: The Stack v3 Released as the Largest Open-Source Code Dataset(3 posts)→
More from Research
- Hugging Face releases The Stack v3, a 5T-token open code dataset — lvwerra · 2026-07-23
- Mila Quebec will host a talk on what AI benchmarks really measure for African languages — hugo_larochelle · 2026-07-23
- An agent stack diagram says production AI is 90% architecture, not prompts — theomitsa · 2026-07-23
- MIT’s free “SLAM for Dummies” guide turns robotics navigation into a hands-on tutorial — lukas_m_ziegler · 2026-07-23
- Sol 5.6 reportedly writes a full research paper from one prompt — conitzer · 2026-07-23
- NeurIPS position paper reviews are now out — hiddenmarkov · 2026-07-23