Hugging Face Releases The Stack v3: Largest Open-Source Code Dataset
Hugging Face has launched The Stack v3, the largest open-source code dataset comprising 5 trillion tokens and 114TB of raw data across 770+ programming languages. The team emphasized the critical need for robust open-source models in the current landscape.
2026-07-23 ~ 2026-07-23 · 4 related posts
- The Stack v3 expands to 114 TB with 224M repos and fixes a dedup bug — anton_lozhkov · 2026-07-23
- The Stack v3 releases 5T code tokens across 700+ programming languages — LoubnaBenAllal1 · 2026-07-23
- The Stack v3 arrives with 5T training tokens and 120TB of raw data — lvwerra · 2026-07-23
- Hugging Face releases The Stack v3, a 5T-token open code dataset — lvwerra · 2026-07-23