Hugging Face unveils The Stack v3 with 5T tokens for open code models
lvwerra · x · 2026-07-23
Hugging Face releases The Stack v3 with 5T tokens and 120TB of raw data
Hugging Face says the new The Stack v3 dataset is being built for open code models and, in the author’s words, for cyber defense needs.
Key details:
- 5T tokens ready to train
- 120TB of raw data
- positioned as the data foundation for future open-source coding models
The post frames the release as part of the broader push to ensure open-source AI remains competitive, especially in security-sensitive coding use cases.
Related event: Hugging Face Releases The Stack v3 Code Dataset(5 posts)→
More from Research
- SaTML 2027 adds first-ever competition and workshop tracks, deadlines set for Aug. 28 — thegautamkamath · 2026-07-23
- Harness Handbook maps agent behaviors back to source code and lifts plan accuracy — omarsar0 · 2026-07-23
- Low-light RGB SLAM holds up only with inertial fusion and global optimization — ucu-autonomous-ugv · 2026-07-23
- GEN-1 now supports everything from five-finger hands to specialized robot tools — k7agar · 2026-07-23
- Google maps 15 million Gemini interactions across 800 jobs and 140 languages — sebkrier · 2026-07-23
- Fields Medalist Tests ChatGPT: Solves PhD-Level Math Proofs in Under 2 Hours — 新智元 · 2026-07-23