Hugging Face releases The Stack v3, a 5T-token open code dataset
lvwerra · x · 2026-07-23
Hugging Face releases The Stack v3, a 5T-token open code dataset
Hugging Face Research says the new dataset is meant to help train the next generation of open code models, especially for cyber defense.
- Scale: 5 trillion tokens of ready-to-train code data
- Raw corpus: 120TB
- Purpose: build stronger open models for cybersecurity work
The release comes as the team frames open infrastructure as part of the response to recent cyberattack concerns.
Related event: Hugging Face Releases The Stack v3: Largest Open-Source Code Dataset(4 posts)→
More from Infra
- DeepSeek-V4 post-training on Ascend SuperPOD reaches 34.22% MFU — _akhaliq · 2026-07-23
- Alphabet’s free cash flow goes negative as AI capex keeps climbing — firstadopter · 2026-07-23
- Huawei compute limits leave 800B-model training far out of reach, industry source says — ShakeelHashim · 2026-07-23
- Atomic says its quantized-model runtime cuts KV cache use by up to 6.4x — testingcatalog · 2026-07-23
- DeepSeek says it has about 20,000 H-equivalent compute cards and will keep buying NVIDIA GPUs — ShakeelHashim · 2026-07-23
- Can an M5 MacBook Pro with 24GB or 32GB RAM run Qwen 3.6 27B locally? — Adventurous-Gold6413 · 2026-07-23