Hugging Face unveils The Stack v3 with 5T tokens for open code models
lvwerra · x · 2026-07-23
Hugging Face releases The Stack v3 with 5T tokens and 120TB of raw data
Hugging Face says the new The Stack v3 dataset is being built for open code models and, in the author’s words, for cyber defense needs.
Key details:
- 5T tokens ready to train
- 120TB of raw data
- positioned as the data foundation for future open-source coding models
The post frames the release as part of the broader push to ensure open-source AI remains competitive, especially in security-sensitive coding use cases.
Related event: Hugging Face Releases The Stack v3: 114TB Largest Open Source Code Dataset(13 posts)→
More from Research
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Sample selection and ordering matter a lot in LLM training: DataFlex makes data scheduling dynamic — Puzzleheaded_Box2842 · 2026-09-11
- Jeff Heaton's Intro to the Math of Neural Networks eBook Is Free to Download — blaizedsouza · 2026-09-11
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11